Why and how COBOL is still used
medium.com
medium.com
The whole stack, and the complex business built around it over many years makes it so difficult for these large institutions to port their solutions on to a new platform in my experience too.
For a lot of these companies, these backend systems on Mainframes just works and I believe there isn't a proper incentive to move on to a new platform anyway.
COBOL itself isn't too difficult to pick up, its the rest of the infrastructure around it. JCL is particularly unforgiving. I'm in total agreement with the parent post of this thread.
I think the thing that bugged me the most about COBOL is the difference between END-IF. and END-IF
Everything else was onerous compared to other languages but relatively straightforward. It's not something to get a simple job done quickly in though.
The simplicity created by things largely remaining the same over time in mainframe world is one of its core strengths.
Curious if anyone has been tasked with replacing COBOL applications and what you eventually used to do so? Being of the web-programming persuasion, I was thinking Electron or Reactive Native for Windows to build something where I could have more control over the navigation experience and then focus heavily on keyboard navigation in the app.
The business likes more interactivity, more visual UIs, but sorely misses what I mentioned above COBOL (or terminal-style) applications in general. Curious where a happy medium might land in terms of language/framework.
I think, as you stated, we just have to double down on layouts, tab progression, etc and rethink what we think we know about web applications.
(I've worked at a security related startup that worked with mainframes which used cobol... The security practices of the runtime do not seem to me less advanced than what we do in other languages.)
Over half of the code I worked with was input validation (format validity, value plausibility, authorisations).
Most of the rest was post-facto batch checks for the same sorts of things, except more.
Any CRUD code base that has been in use for a while is mostly this kind of thing.
I don't think there are any SQL-style grant/revoke directives within COBOL, and mainframe developers in decades past likely did not spend too much time worrying about system security when the best I/O was a dialup modem.
And now we attach this old code to the internet.
That is not a COBOL feature, it's a operating system level feature. IBM System/390 and newer store data in a indexed, record-based manner.
> I don't think there are any SQL-style grant/revoke directives within COBOL
I don't think there are any SQL-style grant/revoke directives within any language that is not SQL, Ada or Rust.
To be fair, code dealing with grant/revoke should be done in migrations, not in the application code. If application code can do SQL grant/revokes by itself you have bigger issues anyways.
It's true that MF tech stack (z/OS, COBOL, CICS, DB2, ..) has many warts and is not fashionable. But I think you can write applications for it in the modern way if you want to; I think the backward-compatibility allowed IBM to focus on continually improving the stack rather than chasing latest fashion. As someone quipped, these little COBOL programs communicating in CICS MRO do resemble microservices or lambdas, and JCL/RDO is just infrastructure as code. So many of the "modern" concepts have been there for decades now, and they let you do transactions very reliably.
But I want to point out, interestingly, what gives us real headache are the modern applications that are NOT COBOL. They tend to call databases with dynamic SQL (basically the plan is created during execution), and this makes the undesired application changes quite difficult to trace. In old-fashioned COBOL applications, you compile your query plans with your application, and that makes the performance more predictable.
So I wouldn't be surprised if traceability is a big thing that COBOL and MFs have still going for it. There seems to be more middleware tooling for that, and things have been standard for many years. Unix tools seem to be more fragmented, although I have no doubts there is lot of good commercial tools. In cloud, everything seems to be even more tied to a single vendor than mainframes - I am not even sure if it's possible to run 3rd party utility on your deployment. It looks like if you really want traceability you have to do it yourself.
Overall, I think MF stack is very decent (although not so shiny on the outside) and I wish IBM would promote it more.
I wonder if the same is going to happen with newer companies, the software reinvented again. Had we have copyright lasting only for 20 years (like patents), it would be possible to legally run first version of z/OS today.
… hold my beer
Having the right tool for the job is always a joy.
Is it just that there is so much COBOL and it’s not worth the effort?
Edit : thanks for all the great replies - HN at its best!
In my experience those things not always go together.
I've seen "mission critical" software with basically zero formal specification.
Maybe this also happens with COBOL?
Didn't quite happen in COBOL's case.
The little document that did exist was hopelessly out of date. A lot of business was done on that software
COBOL is very easy to learn, but, as was pointed out in [1], it's not only COBOL, but COBOL and CICS. CICS is an extremely optimized transaction manager that's been continuously improved and has been coevolving with IBM hardware since the 60's.
There is also the workflows around the very non-Unixy nature of zOS. IBM has been doing an excellent job at introducing good operations practices and adding tools and functionality to facilitate that, but it's a long process.
Also, remember these machines also run Linux incredibly well. They are unbelievably fast.
> If the systems are so vital surely there is a detailed specification
This is hilarious, and wrong. There probably is a specification somewhere, but there is drift between the implementation and the specification. When that 4am bug pops up and stops processing a billion dollars worth of transactions because of a weird corner case, someone fixed that to get the system running again, did they update the master specification, maybe. Half of the job is figuring out how the system really works, all the bugs, all the quirks, all the things that other systems have come to expect.
Sunsetting is often a bug-for-bug type of conversion, because the output of COBOL programs are similar to database tables. Other programs are built on top of this data and that consumer code introduces a bunch of expectations on how the output of your replacement must behave to keep everything working. The number of consumer programs that have some expectation on the data can range from dozens to thousands.
So you can either read and understand and fix all the consumers as you make your replacement behave to spec instead of behaving identically, or you do what most sunset teams do and you make it produce identical output.
This now means that you are reading the target code base and understanding hundreds of thousands of lines of COBOL, you are unearthing decades of extensions, bug fixes, hacks, and mistakes.
Your job then is to make sure the replacement behaves exactly the same way over decades of input. This is the best test of course, learn JCL and beg the mainframe operators to let you run a massive job that will process historical records with the target and replacement programs and compare the results.
Your days begin to look like the following. Get in, get the report from last nights run which tells you all the discrepencies. For each one you can run your code and the cobol locally with some work. Formulate a theory for why some figure is wrong, fix it, move onto the next. Do that for 8 hours. At the end of the day kick off your massive JCL and do it all again tomorrow.
Just do that everyday for a few weeks or months or however long it takes until your replacement program behaves the same as the COBOL program and then you can swap them without breaking 1000 other applications.
Then you get the long tail of bug reports. The zipjack flim-florp report isn’t working. So now you dig into that, the flim-florps are processed by the COBOL you replaced, so maybe the zipjack system had a requirement you didn’t realize. Dig dig dig, days go by, your will to live eroding, but finally you find it, deep in the CVS history bug-1376892 looks like the target program accidentally added an extra hour to the business day in Seoul on leap years because the time zone data is wonky and they have special code expecting this wonkiness and correcting on their side because the mainframe programmers in 1994 decided it was too much work to fix on the mainframe side.
Our program uses a different timezone database and it knows how to properly calculate this, so just add a data transform at the end to detect this scenario and fudge the numbers back to what the zipjack (and who knows how many other systems) expects.
The COBOL systems are rarely used directly, they are batch processors taking in data and pumping out data. Applications are then written on top of their output and bind the behavior of the COBOL system with all their expectations. And this is where the complexity lies, the specification for how the system works is possibly written down somewhere, but the actual specification is in the COBOL code and the hundreds or thousands of consumer programs and all their expectations. To do the job well you have to write something that replaces the actual real-world specification encoded in all that other software, not the one jotted down in wordstar a few decades ago.
I think that the missing element in my initial analysis is that many firms don’t have much in-house expertise left on these systems. They’ve cut to the bone or outsourced.
And for management replacing is huge risk for limited upside - even if the economics make sense.
> So you can either read and understand and fix all the consumers as you make your replacement behave to spec instead of behaving identically, or you do what most sunset teams do and you make it produce identical output.
So on point. And this problem exists not just in legacy systems built in COBOL. Also in much newer systems built with SQL in my line of work.
The horror show I regularly encounter is a SQL function used in a production system that requires 50+ variables as input. 50+ fields as output. Any change in the order of the fields screws everything up. Makes me crazy.
For testing you'd run a few weeks of data through and make sure they match outputs exactly. In prod we'd run them in parallel a bit for each report before decommissioning it.
There'd always be something like..at the end of 3rd quarter if the last day is a Friday, some other junk happened and now your data looks different.
After a few we got a hang of which things were likely to break and was able to have certain parts driven by configuration and many could be fixed via an updating a config if a bug was found.
If you've ever worked a VisualBasic Application, you'll have some inkling of what COBOL is like.
One of the core things I saw with COBOl was strong coupling of application and data persistence. Further, where modern apps take a stateless approach for say a single API call, many COBOL application predate this concept. So not only are application and persistence coupled, the applications are VERY stateful.
This makes replacing a 'piece' of the application extremely difficult as there is a lot of shared state to tease out of the system. Nor can you start building a second system that owns some piece of the data model, as COBOL's persistence models aren't built for shared write activity that a normal DB facilitates. You are forced to still write through the COBOl data model.
Clearly not impossible, just difficult.
Many of these old systems rarely require changes in most modules. (As in, of the hundreds of thousands of lines in the system, many subsections just never get touched).
This means for the the parts that never get touched, there is a negative return on investment for rewriting them, so there has to be a correspondingly larger positive return on investment in the higher churn parts to make a rewrite worthwhile. Most of the times when this is studied, the return on investment from the rewrite is either tiny or net negative from this effect.
Ah you say, then just rewrite only the parts that tend to change a lot, and leave the other stuff as COBOL. That sounds good right? Except that generally the cross language interop code needed becomes an absolute nightmare. I've dealt with this in an near ideal enviroment, and it was still a royal pain.
This near ideal environment was: COBOL code restricted to just a subset used by the company, which was being translated to C code by the companies own compiler, which allowed for a C-compatible ABI to be defined allowing for surprisingly simple and direct C to Cobol interop. Yet even then the code code was still a royal pain, just due to how COBOL data structures differ so much from the idiomatic data structures for C or C++ code. My code was actually C# code interacting via C++/CLI with some C++ code that interacted with the the COBOL. But the interop code was still some of the worst code to touch.
> Realistically If there were detailed specifications, they were hard copy print-outs from the 80s that are probably no longer readable, and undoubtedly were thrown away in the intervening decades.
Or not kept up to date. My employer has some mainframe systems, and when they were built they had technical writers and created pretty good documentation, organized in binders.
For the last two decades, at least, these systems have been "legacy" and the company stopped employing technical writers. During that time, they were still pretty critical, but I'm pretty sure no one kept that documentation up to date with whatever changes were made.
Also, about a year before the pandemic, we had building reorganization, and a lot of those binders were thrown out during the move. They were kept on several large, neglected shared bookshelves, and I think anything anyone didn't specifically claim was discarded by facilities.
The thing with legacy systems is they can be vital, but they also can be neglected. The gamble most businesses seem to take is: "as long as it works, starve it and hope it doesn't blow up on my watch."
Too many people learn this specific COBOL fact the hard way.
https://medium.com/the-technical-archaeologist/is-cobol-hold...
Also consider that the COBOL datatypes are directly represented in the underlying storage, so whatever format you decide to represent a value in in the code is the same format that will be used for external communication.
If you want to retain compatibility with COBOL, you're probably going to need a COBOL-compatible numeric library in whatever language you're using. at least if you want to just reimplement the original code.
It was standardized two decades before ieee-754, that we now take for granted.
The math is wildly different.
As others have said, no, there's probably not a complete specification. But even if there is, it's not nearly complete enough.
That code may have 50 years of bug fixes and corner case fixes in it. How many times was the spec changed to reflect those fixes? Probably not always. If you rewrite it from the spec, how many of those issues are you going to re-introduce?
And, how many new bugs are you going to introduce? The old code has 50 years of "getting the bugs out". It may be hard to maintain, but for what it does, it probably does it pretty solidly. Will the new code be as solid?
Re-writing isn't as easy as it looks...
That said, all that code has been running for 3 or 4 decades and anything that gets ported will have a shelf life of < 10 years.
The author surely realises that there's this thing called LPAR and z/VM? And that you don't just run a "single COBOL application" on one of these boxes?
(Source: I met one of the debuggers while I was in Patagonia. He is in his 70s and sails around the world with his wife and fixes bugs when he's near WiFi. Either he was telling the truth, or he's a really good storyteller.)
Making it 1000x worse than something that would already cause me headaches and nightmares.
https://www.wealthsimple.com/en-ca/magazine/cobol-controls-y...