How popular is COBOL?
techradar.com
techradar.com
I ended up going to neither and went to a traditional CS undergrad at a 4 year state school, and things have been fine ever since. However, on a regular basis, I think about the other life I would have had, if instead I had gotten an associate's in mainframe programming. What would I have done? Where would I have worked? Who even are those folks.
If today I found a COBOL bootcamp that reskilled developers for a career change into COBOL, i may even seriously consider it. An unexplored path that has always been in the back of my head, deep in the guts of the engines of industry.
Cobol devs were on average a bit older (on average id say 10-15 older than the rest of the devs), mainly because there weren't many college grads going in, and like your advisor had said there was a split between cobol and non-cobol devs with very little overlap.
The majority of the cobol devs seemed to fall into a few camps. There were people who graduated in the 80s or 90s and just kept doing the same thing. There was also a large amount of Eastern Europeans that had immigrated in the 90s / 2000s and were cobol devs. Then there were also a decent amount of Indian contracting companies (TCS and the like) which seem to also train people in Cobol even today to hire as contractors.
It's like an alien world where some of the ideas are similar but also everything is completely different. Linux/Mac/Windows are all different but they are not different like working with mainframes is different. So if you started without any PC background, it really doesn't matter.
Personally, it wasn't for me. The whole thing was basically frustrating. The technology was frustrating. The glacial pace of development was frustrating. One program I worked on was literally older than I was. The code quality was very high (given the technology) but that was accomplished by sheer effort rather than having smart tools.
https://blog.cloudflare.com/cloudflare-workers-now-support-c...
What attracted me to it is just how different it is from the UNIX/PC world I'm coming from, it's endlessly interesting, it challenges what you assumed were ground truths. So is AS/400 (or IBM i as it is now called), which is also completely different but in some ways opposite to how the mainframe world is different.
Those two rewired my understanding of OS reality like Haskell rewired my understanding of programming language reality.
At this point I could probably dive into COBOL and pivot my job into maintaining mainframe stuff, and some part of me thinks it's fun, but I have a fun low level programming job, so I don't know if I'd ever really do that. Maybe to fill time during retirement?
So to be clear, what I meant to say was that for a small hobby I dived relatively deeply into it (i.e. I see others messing with it for fun not caring that much about e.g. JCL), and that I wet my feet enough that if I put the serious work in, I could see myself eventually becoming productive in the COBOL and mainframe world.
Just because rereading it, it sounded like I said I became mainframe proficient to a commercial level by playing around with a many decades old MVS version on Hercules, which is hilarious.
If you have a solid background in C, it is not very hard to learn. It seems to show up at older companies, and is fairly good at moving and reformatting records where you want to address each field on a character by character basis.
Unlike more modern languages, there's no encapsulation, polymorphism, local variables, etc., so you can write tangled code quite easily, though the code I've seen has been very well written. Everything is in capital letters, and the character set is often ebcdic. Most numbers are fixed decimals, which leads to some unusual math results.
As a result, it can be maddeningly verbose -- I've seen 100+ lines of code that would just be a single printf statement in most other languages.
I want extra money to code in a terrible language maintaining your bloating, ancient codebase. They believe that all of this should mean LESS salary.
Meanwhile, old COBOL devs are retiring in droves and the businesses stuck on COBOL have countless billions riding on those systems.
Why hasn't supply/demand kicked in yet?
Therefore India's big IT shops can simply train some fresh grads to learn COBOL. Those grads don't care about working on the latest tech, they just want a well paid and stable job. This puts a hard ceiling on the wages of COBOL programmers.
The only development environment they know (barely) is Excel
it’s all about saying no to raises, controlling people to work with their butts in a chair in an office, 8am meetings, 10 different teams to do basic things, 5 people per project whose sole job is to say “no we can’t do that”, lots of 2am batch processing on call
not sure if supply/demand works for 40k employee companies
I understand that it's not ideal from a language point of view, but does it have any properties that make it a great fit for the job? Does COBOL on a mainframe have particular properties that make it good for concurrency, validation, etc?
But mostly, it's just impossibly expensive to rewrite a humongous spaghetti of societally critical systems from cobol to something modern such as java or .NET.
When I think about DAGs today, i think about Airflow. And SQL.
Would it be a better match to rewrite these systems in SQL and Airflow? SQL for the logic and Airflow for the batch processing.
I know for many (particularly those who mention java and .NET) SQL is just a place where you fetch and store your data. But once you start building systems with it, you will soon realize it contains nothing more and - at the same time - nothing less than what you actually need for mangling your data in a terse way.
I know there are many reasons to frown on SQL for this, and I am fine with any comments about it. I think it can be a start of a good discussion. Nothing is black or white.
All the actual customer data is stored in databases so that's why SQL is needed. And then to access those databases you call various in-house cobol systems with the right parameters. SQL is good for the type of business logic that these systems do, for example niche cases like FATCA and CRS tax reporting and tons of legal details like that, to organize the business requirements it's embedded in the cobol programs as DB2. When you want to send out mails daily, monthly, quarterly, yearly, and so on, in ten different modes depending on a parameter, with different variants depending on user and organization data that's an orchestration done in mainframe OPC scheduler with various applications with jcl and proclib, and I think SQL is too strictly logical and can't do the fuzzy parts with strings and files and system communications to be economical as a full replacement. I'm pretty new so I don't understand most of the stuff honestly which is why I'm throwing out a word salad here.
We do have java and .NET parts that we communicate with, built around bought solutions that needed integration. Airflow would have to be integrated in the same way but can't be a full replacement
Handling the DAG architecture in a batched system is exactly what airflow does. Do this, wait here, when it is done do that. So the DAG acronym is entirely appropriate for this discussion.
And you can have tasks in Airflow that send out mails, quarterly that depends on the quarterly summary to be executed.
So my idea here is to
- Ingest all raw data into e.g. BigQuery. - Combine it the way you want with SQL - Add tables with email addresses etc for customer adaptations along with parameters - Join with those table to create custom adaptations. - Add output layer stuff (email, upload to custom file system etc).
Now you have one place and one language for your logic. Given that this part of the system is inherently batched that is. If it is real time/streaming it will not work.
The "can't do the fuzzy parts with strings and files and system communications" needs to be defined. This is the rot in the system, the ghost in the closet that everyone is afraid of approaching.
It starts off by suggesting that it's all about how COBOL's superiority stems from having support for Binary Code Decimal as a language-level element, rather than having to be imported via a library (the overhead of which really starts to matter at the volume of transactions which COBOL is typically required to handle). But then broadens the discussion out to argue that the intrinsic shape of the COBOL environment "stack allocation, pointers, unions with no run-time cost of conversion between types, and no run-time dispatching or type inference" is fundamentally different from languages like Java or C#, and those differences provide provide performance benefits which cannot be easily obtained in those other languages.
https://github.com/wmealing/cob-webber
Ncurses with vt100 emulation kinda works. But it's timing is garbage and text input is funny.
As a way of learning the language I end up writing a perceptron in it [1]. It was a fun exercise.
Telcos, healthcare providers, insurance companies, anything that was early to get a billing system.
Don't get into COBOL if you get annoyed at bad code :D
By all means, believe whatever you like about Cobol's popularity. But if you think TIOBE has any weight, reconsider your sources.
Rather than asking for a better alternative for gauging "popularity", stop looking to aggregators and instead look at actual numbers. For example, if you want to know which languages have the most job openings available, don't trust a measure of popularity, go to a job board and search by language.
However I have a deep distaste for writing in COBOL. That said, my favorite resource for dealing with COBOL problems was AMBLIST which shows the compile date and other attributes of each object module,whether COBOL, FORTRAN, Assembler or PL/I.
There's a number of mainframe source management tools. In VM, XEDIT would automatically generate change records for source files as they were modified.
It made a lot of sense to develop under VM and deploy to MVS, z/OS just to have good source control without having to buy source management tools and enforce their use. I strongly suspect IBM does pretty much the same for their mainframe operating system and utility code.
The dirty secret is that the code that runs logical partitions is a flavor of VM.
COBOL and FORTRAN in the early days had minimal subroutine libraries. Direct calls to the supervisor were made in the compiled code along with building the parameters and handling the results. This does make it more straightforward to see what the object code is doing.
PL/I had a comprehensive subroutine library.
Later on the subroutine libraries for the various compiler languages were consolidated into Language Environment.
I find it hard to believe banks are running critical code they didn't at least back up somewhere.
I also don't understand why it would be hard to back up the code.
If that sort of development can't force a rewrite then nothing will.
Sometimes the COBOL isn’t the problem, it’s the scaffolding built around it. As an example, I worked in a consulting role on stabilizing systems during COVID that were getting crushed. The COBOL stuff on the mainframe wasn’t a problem - tweak the job control stuff, give IBM more money to lease more cpu time, and you’re good to go.
The problem is the 1995-2005 legacy middleware dreck that isn’t engineered well like a mainframe app and doesn’t scale. That stuff is always hot garbage and makes it difficult to replace the backend mainframe. The middleware expects whatever comes out of the mainframe, you end him with a chicken and egg scenario.
Financial services tend to have good data management practices and regulators who know how to audit for them. As a result are successful in getting off mainframes. Governments, manufacturers, etc, often struggle with that.
If that happened in the 80s or 90s the code that exists has probably diverged enough that it's effectively pseudocode at this point.
(I know IRS has been working on rewriting it all - the 7074 assembly, the 360/370 assembly, and the COBOL - in newer languages such as Java. I don’t know where they are up to with that, so I don’t know if they still have 7074 apps live today. But certainly they did as recently as 4-5 years ago.)
I've witnessed second system effect.
Then, they tried again with a competing system from a local Australian ERP firm (TechnologyOne), running on a Microsoft application stack and Oracle RDBMS; and this time they succeeded in replacing the mainframe COBOL system, in part because (unlike PeopleSoft) it was actually designed for the Australian market.
[0] https://www.fujitsu.com/au/about/corporate/history/products/...
I may have been unclear, but this was the scenario I was describing. The need for a bug-for-bug re-implementation is a special case where the acceptable deviation from the original program is quite low.
The simple fact is that they've spent tons of time slowly coding in hundreds of thousands of little, undocumented edge cases all over the place. The rewrite gets 90% there then falls flat with all the little things and you can't switch these kinds of businesses until it is 100% there.
Is it not possible to find a safe subset of cases that can be handled by a new codebase. You use shadow testing to help prove the rewrite handles the cases the same as the old implementation and then grow that subset until it encompasses the whole thing. As the old implementation is used less and less you can run it on a smaller computer and have to maintain it less.
This process also outlasts aging managers. Chances are that leadership will shift at least once and decide it’s taking too long So they kill the replacement, switch languages and/or teams, and try yet again.
I know the tooling for doing that doesn't currently exist... but paying someone to develop it would be cheaper than any one of these ground-up rewrite projects!
The problem is that a program doesn't actually contain a full description of the business problem - it only contains what's necessary for the machine to solve the business problem. For example, the system may never be designed to process check deposits on a Sunday (and a lot of date-related calculations will go wrong if you tried to), but guarding against that is accomplished by the "real world" - there's nobody there to file in check deposit requests on those days. Yet, there's nothing explicit in the code about preventing check deposits on a Sunday, so your hypothetical decompiler will have no way to infer that without outside knowledge.
What you need is a hypothetical magical decompiler that given a binary or source code as input, will spit out the actual business problem and all the edge-cases it's meant to solve. AI might eventually get us very close to that, but ultimately it'll still be a guessing game.
AFAICT the goal of such a translation effort should just be to replace a COBOL codebase that does X, with a codebase in a more-maintainable language that also does exactly X, in a bug-for-bug-compatible way. You don't need to know the business requirements that led to the implementation; you just need to take the existing implementation, with its embodied institutional knowledge, and blindly preserve that knowledge by not making any top-down assumptions about the requirements that would accidentally design it out, but instead preserving the bottom-up structure.
If you like, you can re-do the requirements analysis and top-down refactor the program later on — after you've already "decompiled" it into a workable HLL.
An analogy. Say there was an unfinished manuscript of a novel written in Old English. You want to have this published as a modern novel. You have two choices: either you — someone who presumably can read Old English — try to do a top-down rewrite of the manuscript into a modern-English novel, figuring out the plot and characters from the manuscript and copying that "design" to a new modern-English novel; or, you just translate the manuscript itself into modern English, sentence by sentence, and then, after you have now acquired a modern-English manuscript, set about turning that manuscript into a finished novel through the normal prose editing tools of structural editing and copyediting.
I think you can see how the former approach is pretty silly; and yet is exactly what people do when they attempt to build a new system from the ground up to replace a legacy system. The former approach severely limits the number of people who can understand+reference the old design when working on the new design, while the latter approach does not. The latter approach allows the "turns of phrase" of the old work to be reused as-is in the new work, while the former approach does not.
They presumably do this for market segmentation reasons, however if you account for the cost of 1) maintaining the old COBOL/mainframe system (as it's the only thing that's proven to successfully run those products) or 2) successfully moving all that business logic & edge-cases to a modern system, it could emerge that maintaining those different products costs more to operate/maintain that what it brings in terms of revenue from the market segmentation approach, and they're better off simplifying their product line. This also translates to other areas of the business - less support overhead, etc.
The other problem with ordering (and delivering) a successful migration is that a lot of people's jobs in those legacy companies rely on the tech being ancient and would be obsoleted should the tech get better, so there's a lot of internal pushback against this. A lot of people may not actually want the migration to succeed, or at least not to anything actually good as that would put paper-pushers out of a job.
The only way IMO to successfully deliver this as a large company is if the top brass effectively sets up a brand new subsidiary, gives it a budget and a task to make a product competitive with what the parent company offers, and when it's ready they can just move off their customers to that new product operated by the subsidiary. Rinse and repeat until no more legacy tech remains. The fact that it's a completely separate company means it remains lean and doesn't have to be fighting against deadweight that sabotages any modernisation attempts to protect their own jobs.
It's not so easy to "simplify" out features which are legal requirements.
But as others have said, banks and other large companies that started using computers a long time ago would have critical services written in COBOL.
[1]: https://old.reddit.com/r/programming/comments/cd1g2/the_term...
The RPG language PREDATES Cobol! [0] https://en.wikipedia.org/wiki/IBM_RPG
But why would you do that? I suspect there’s few good reasons for doing it (ease of finding new programmers) and many good reasons not to do it (speed, it’s code that has been battle-tested for ages).
Anyway, from friends working in the banking sector, nowadays the mainframe code is exposing an API and users are accessing data via a web interface and not a terminal emulator anymore.