Don't hate COBOL until you've tried it
opensource.com
opensource.com
COBOL is still pretty efficient for tasks like scanning data on magnetic tapes--not that most of us do that any more. In fact it was pretty good for any problem that involved reading from or writing to storage devices because you could control the data layouts very precisely. It was great on reports for the same reason. I even wrote a simple doc markdown processor in COBOL because the place I worked didn't have usable word processors and we needed something to do pagination and formatting of docs.
I don't miss the verbosity or the lack of modularity but on the other hand COBOL was an effective tool that you could use to solve problems more quickly than other tools available at the time.
If languages like Smalltalk, APL and Lisp aren't counted, then sure, Java was big productivity boost in the history of programming.
Einstein got the Nobel for the photoelectric effect, which we take advantage of continuously every day. I'd say that was influential.
I'm also a big FORTRAN fan, especially the VAX-FORTRAN variant DEC introduced in the 1980s. It was great for numerical analysis but also included library calls to access the operating system, much like C. Some of us were late to the Unix party, so this was a huge step forward.
You can do that in assembly too
struct foo {
int count;
char name[16]; /* not null terminated! */
};
REPL: define an alias name foo for the type with typedef: 1> (typedef foo (struct foo (count int) (name (array 16 char))))
#<ffi-type (struct foo (count int) (name (array 16 char)))>
Now put an instance of the Lisp struct into a binary buffer using this FFI type: 2> (ffi-put #S(foo count 42 name "ABCDABCDABCDABCD") (ffi foo))
#b'2a00000041424344 4142434441424344 41424344'
Now, recover a new Lisp struct instance from this binary struct: 3> (ffi-get *2 (ffi foo))
#S(foo count 42 name "ABCDABCDABCDABCD")
No problem; the FFI type system knows that an "array of char" is different from a null terminated string, and can make it correspond to a Lisp string in both directions.Now, for fun, let's poke a zero byte into that buffer:
4> (set [*2 8] 0)
0
5> *2
#b'2a00000041424344 0042434441424344 41424344'
There it is. Now decode: 6> (ffi-get *2 (ffi foo))
#S(foo count 42 name "ABCD\xDC00;BCDABCDABCD")
What's that? My UTF-8 decoder treats the 00 as an invalid byte, and maps it into the surrogate pair range U+DCXX.
The otherwise optional semicolon was output because the next character in the string is a hex digit.If that U+DC00 is encoded back, it will reproduce the null byte:
7> (ffi-put *6 (ffi foo))
#b'2a00000041424344 0042434441424344 41424344'Two variables holding common license plate formats might look like this:
01 LIC-PLATE1 PIC 9A(3)9(3) VALUE '1ABC234'.
01 LIC-PLATE2 PIC 9(3)A(3). VALUE '123ABC'.
[edit: added a newline to separate my examples]
("9" specifies a numerical digit, "A" specifies an alphabetic character, and there are other types that deal with sign, explicit and implicit decimal points, etc).
Assembly deals with memory layouts, but doesn't enforce any kind of higher-level data type, and especially not a human-oriented data specification.
So, most recently I wanted to simply read a file, do some simple processing and then write an output file. Very basic, right?
Not so much with COBOL.
First you have to statically declare the files you want to use. Let's say you want to read from "foo.txt":
environment division.
input-output section.
file-control.
select foo0 assign to "foo.txt"
organization is sequential
access mode is sequential
file status is fs0.
You then need to declare (statically again, of course) the structure of the records in the file. You can think of this as global variables that gets filled in when reading from the file: data division.
file section.
fd foo0.
01 foo0-rec.
02 foo-name pic x(40).
02 foo-value pic x(20).
You also need to declare the "fs0" variable: working-storage section.
01 fs0 pic 99.
Then, you need to open the file. That's not too hard: open input foo0.
Then you want to read from it. You can do that using the command: read foo0.
After the read, the variable "fs0" will be 0 when you've reached the end of the file, so you need a loop that, I presume (I never got this part to work right), would look something like: perform with test before until fs0 = 0
read foo0
* foo-name and foo-value now contains data from the file
end-perform
After all of this, I decided I wanted to do something different. I wanted to call a web service to load some information over HTTP. I dropped that project really quickly once I realised that the only way to do that was to use libcurl and directly access it using FFI. The FFI that is provided by GNU COBOL isn't really easy to work with, as evidenced by this discussion: https://stackoverflow.com/questions/26367026/how-to-make-htt...If you look at the code that is linked from that page, you'll rather quickly realise why COBOL is not such a great idea for general purpose computing.
Outgoing web requests are usually done via either a program in some other language that you call from the COBOL application, or via DB2s built in functionality by running it as an embedded SQL section in your program. Why the heck web requests are handled by the relational database is.. Well I mean it feels odd, but by doing it that way they added the functionality to both zOS and IBM i at once. And old developers don't have to learn all to much new stuff to do it.
Also, while the COBOL program works with a static set of files with static filenames. Those names can point to any file by various overrides done in either JCL on the mainframe or CL on the IBM i/AS400. Regarding the file formats, you can read the file into different working-storage defined variables, so each row doesn't have to be identical field length/position. In short, COBOL is just one part of the solution. Various "scripting" languages play a large part in defining what files go in and out. And none of them are available outside of the actual mainframe/midrange environments. Most data is gathered from relational databases anyway nowdays, and embedded SQL, just like in PHP webpages of yesterdays web, is whats used in newly developed programs.
But all in all, yes it's a rather cumbersome language, for any sort of sane string handling you have to revert to embedded SQL. But most problems you have to write programs to solve can usually be broken down into smaller problems. Then just chain programs to handle it. The language itself is usually just a bit of a bother, but the age old programs written 20-30 years ago written by developers of varying degrees of sanity who had no requirements on them to document their work is what makes the job a pain.
1/ COBOL run in closed environments. This allows these to be rock solid. (no such thing as a dependency nightmare). You don’t have such a thing as different version of Java on your testing and production environment.
2/ COBOL forces you to mix concerns : DB access, file I/O, business logic. This is sometimes useful to get a complete picture of what’s going on in a program. In Java, you have to wander through your DAO, data model, business logic, etc. which are spread everywhere.
3/ Code resue is hard to acheve so there is some copy/paste. That is sometimes good in the sense that if you fix a program, you won’t accidentally introduce a bug in a copy of that program. (many of us will scream at reading this, but consider it).
4/ Of course COBOL programmers are good too and they use these old practices in a sensible way
5/ COBOL technologies don’t change often. So there’s no discussion about changing your framework. When you write software that must last for at least 10 years, this makes choices easier.
Java doesn’t also say you have to separate concerns. You could easily write a single method which mixes business logic and IO (file and DB)
When writing a batch program you need a balance between understanding it, making it efficient and reliable and promoting re-use (e.g. if many batch programs want the same behaviour break it into a reusable module). COBOL still has dependencies too, so each program is rarely written in isolation and you have to co-ordinate with other teams to ensure their module versions are in sync with yours.
The thing that has always let COBOL down for me is the lack of automated testing.
Sure, I was talking from inside my own enterprise-context where Java actually means Java + Spring + a few other things that force you to separate concerns (for the better imho, although it doesn't help with code navigation)
>>> The thing that has always let COBOL down for me is the lack of automated testing.
True.
Java doesn't force you. You can structure your code in many ways. In particular nothing is forcing you to have objects like Customer, Order, Invoice, etc. It is just the way OOP has been traditionally taught. Specifically in enterprise applications where a lot of the logic revolves around the data tables makes sense to have objects like CustomerTable, OrderTable, InvoiceTable and have all the related code encapsulated. See https://martinfowler.com/eaaCatalog/tableDataGateway.html https://martinfowler.com/eaaCatalog/transactionScript.html
All your points make a lot of sense. But it is not an essential part of the language. It is more about the culture prevalent with people using the language. But your point is well taken. Maybe we can learn from other programming cultures.
Both Go and Java can do that (yes, without having to install a JVM on the target machine).
So she signed up with Alcor, and had herself cryonically frozen in 1999, leaving behind explicit instructions that she was to be revived after civilization had finally put itself back together again.
Finally she woke up, in a clean futuristic hospital room, surrounded by inscrutable machines that go "ping" and strangely dressed doctors with funny accents and weird hairdoos.
She was so happy to be alive that she proclaimed "Thank you so much for bringing me back to life! I am eternally grateful, and in your debt! Is there anything I can do to repay you? And by the way, what year is it?"
One of the doctors smiled at her and said, "Yes, actually. It's the year 9999, and the records indicate that you're a COBOL programmer..."
"Dragged" as in it rains money on them until they stop saying "no".
Today is the last of a 7 week-long "Mainframe Bootcamp" during which I've been introduced to COBOL as well as TSO, JCL, DB2, CICS and MQ. Thus far, I'd say the class has picked up the language fairly quickly, however we've yet to fully grasp the intricacies of development using our company's heavily customized mainframe tools. Generally mundane processes, such as compilation and promotion between environments have been, IMO, rendered tantalizingly complex. It's all really quite tedious and boring. IBM's documentation monopoly hasn't helped, either.
To the point - I question whether COBOL's seemingly polarizing reputation can be somewhat attributed to the environment in which it's usually developed/deployed more-so than the semantic's of the language itself.
Of course retrospectively with gigabytes at our disposal it is hard to put yourself in that frame of mind.
Maybe it helps to understand this better if you realize that the assembler that I cut my teeth on only supported 6 character labels and symbols.
But how complex can these programs be? Surely anything written before 1992-ish can’t involve that much code just because of the limitation of the machines they were designed to run on. What makes it that expensive that no one wants to rewrite them in a modern language?
You would be utterly shocked.
I did some COBOL work for a bank.
Most of the new code they write is in C++.
However, COBOL was still a huge part of their stack, and they still had several heavy duty mainframes that it ran on, and maintaining that was part of my job.
For an example of how much COBOL was written back then, and how huge the technical debt to replace it would be:
* The interbank transfer system alone was ~750,000 lines of COBOL.
* It handled around 1.2million transactions every three seconds. That's not capacity, that's real world data.
* They had less than five minutes of downtime in the last twelve months, and that was considered poor.
It's a mammoth task to replace the system, and getting transactions from the old to the new without incurring a huge overhead is nigh impossible. And yes, dropped transactions can mean dropped customers, and big customers.
It's far cheaper to hire and train programmers in COBOL, than to rewrite everything. COBOL and Fortran have millions of lines of code in use, and such frequent use that replacing them may cost more than the lifetime of returns would give.
Sounds like jumping from the frying pan into the fire.
Yep, that. I did a two-year stint in a major financial network. I guess it's not a lot of time, but while I was there, there was a big outage and it was not on the mainframe side of things- some all modern server farm bought the, er, farm.
From what I can tell, at least in that company, mainframes are still in use because their stability is simply unmatched by anything in the "distributed" world; Cobol is still in use because mainframes are still in use; and nobody dares replace the Cobol code with something else, because the code itself is damn stable, as a result of hundreds of people having worked on it for several decades now- and counting.
In general, Cobol on mainframes is a rock-solid system that doesn't fall down and that nobody wants to disturb.
Also, I've yet to see such businesses opt for the safest porting approach. Build a transaction replication system, then a component of the ported system. Inbound transactions are replicated and dispatched to the legacy and the component of the ported system simultaneously. Build scaffolding to monitor key transaction metrics. Compare output and performance results. Let that "bake" for a year or two. Rinse and repeat with the next component (many such ported components can be developed in parallel, etc.), and so on, until the ported system reaches feature parity with the legacy system, then let the completed version "burn in" for 2-3 of the longest periodic process lifecycles. For example, if the the longest periodic process the application supports is once every 7 years, then you burn in for up to 21 years, comparing during the entire duration.
By the time you "throw the switch" to the ported system, you have a very high certainty that it will work just as well if not better than the legacy system. Extremely low-risk, but extremely expensive, and considered extremely impractical. And with how business-critical some of these systems are, pretty much any risk mitigation short of this will get nixed by senior management.
This doesn't even begin to get into other factors, like changing regulatory and business environment as you are porting (so you're porting and modifying for new requirements simultaneously), such a project can easily be targeted for the cutting block for just the cost savings to make someone in Finance look good, internal political issues, developer churn over such a long period, developers on the legacy system withholding key implementation or even business requirements information, that information simply not being available, binary-only parts of legacy, etc.
I would disagree with the assessment that management is too dysfunctional, it depends. But if they all were like you described, i would be out of work. And it just happens that we turned off a mainframe just 2 weeks ago.
From seeing how the risk is assessed in some of my customers, the kind of emulation approach you described (I've seen one for the old HP 3000, and it worked great) sells easily when the legacy vendor is so sclerotic and rapacious that the customer's management is practically pushed out of the relationship, and the economics don't have to completely pencil out to close the sale.
The customer is tied down to a smaller (usually) vendor's emulation, and trust the fidelity of that emulation over time, and furthermore, over new versions. The customer must maintain a skill inventory that isn't "current" (though some people might consider that a plus).
I can think of a couple ways to address those sales objections (lock-in isn't real: can always switch back to legacy vendor if all else fails, no need to revamp staff with crucial business/domain knowledge, etc.). But I'd be really interested in how your sales team deflects them, because I anticipate this legacy software issue will only get worse in our industry over time, and not just on mainframes. I'd like to try to figure out how to bring down the friction of such sales in the future.
Thanks for your feedback, really appreciated.
Aw. I hope you didn't hurt it.
I don't think that's the case. In the place I worked (major financial network) the average time people seemed to have ben with the company was 10 years. Especially in big companies, you can expect to see lots of people who go for a long-term career, instead of hopping from job to job (and promotion).
I think the real reason why management doesn't dare to make big changes is that they understand there's a very real risk to cause a significant loss to the company (and lose their job, as a consequence).
Which is why Cobol shops have absolutely mad processes. Where I worked, to make a change to a Cobol program -even if it was just the JCL that would run it- we had to complete a ten-page form, submit it and wait. They basically just don't let you touch the code. You work around it, sorta.
What do you think about the .NET and JVM Cobol compilers?
They seem to be a way to move that code into new systems with no rewrites as well.
But for the generic case, the mainframe is quite an integrated system, and big clients tend to use a lot of the features at the same time. Even if you can produce an executable or a library, but you ll still need a transaction manager (CICS/IMS-TM), a database (DB2/IMS), MQ, RACF, Datasets, GDG, REXX/JCL, etc. You need the whole ecosystem. And all the binaries will expect a DB2, or a MQ listening on the other side, with near-perfect emulation (return codes, db behaviour such as EBCDIC sorting, etc). Just compiling wont help you much unless you have a very basic use of mostly JCL and batches.
1. Introduction of errors (both overt and subtle)
2. Fixing of what seem to be issues in the original code that are actually undocumented workarounds (or similar things) to potential bugs (in COBOL or who-knows-what-else-in-the-codebase)
Both are big issues in such a move; software engineers don't often appreciate that much of this old stuff is still around because it WORKS. It has been virtually fully debugged, just through the iteration of time and experience.
So porting can easily introduce bugs that may not be apparent immediately; and because software engineers are always wanting to make thing better, if they are working with the COBOL codebase and see something that they think could be done "better" - that change might actually cause problems down the road, perhaps in ways that aren't recognized immediately, or only in edge cases - things that were worked around decades ago in the COBOL implementation.
An ex-colleague used to work for a major US insurance company, and this is what he told me: They had a bunch of products, and due to the state-by-state nature of US insurance regulation, they had a slightly different version of each product for each of the 50 US states. And then, for every one of those different versions of those products, they had a separate suite of COBOL programs to manage them. And, all these COBOL programs basically did the same things, but all with slight variations due to the different rules of each state's insurance regulators. So they basically had hundreds, maybe even thousands, of separate COBOL programs to maintain, a mess of horribly repetitive code. Their problem with moving off COBOL was not that any of the code was particularly complex, just that there was so much of it, and trying to rewrite it into another language (or, maybe more sensibly, translate it into some sort of business rules engine or insurance policy automation product) would take major effort and run a big risk of getting something wrong in the translation. (He worked there more than 15 years ago, so things at this insurance company might be very different by now.)
COBOL will outlive us all.
[Edit: changed period to question mark]
Regarding "how complex can these programs be", these programs run core operations for banks and insurance companies, so you can probably imagine the amount of hidden business logic and technical hacks that have stacked up over decades of operation and changes. This will not surface in requirements gathering meetings without A LOT of blood, sweat and tears. The alternative would be to reenginer the company's processes before the building the software to support them, but that would mean a lot of change and that tends to be harder than most software engineering endeavours.
Not only are these programs complex (projects to re-implement them from scratch in J2EE would certainly be measured in years), but they are mission critical and the organization cannot afford to live without them or something 100% equivalent.
There was a lengthy discussion about this on HN some time ago: https://news.ycombinator.com/item?id=12096250
You'd be surprised. The business rules and exactitude of results can be documented in 1000s of pages.
It’s in a local maximum. Just like all technical debt. And as far as I hear, cobol is used for things that put that risk and liability factor on steroids. I know I wouldn’t touch it without some serious budget to set up massive amounts of redundancy and fail safes.
From a business perspective there may not be much to gain from doing so. If the old software works in a current environment then there often isn't a justification for the cost of rewriting/testing/integrating just to end-up with the same thing in a shiny new language.
I'm happy not to work with JCL, COBOL, EBCDIC, and mainframes in general. I realize just how much awesomeness in my OSs that I was taking for granted.
Want to declare a variable? You're gonna have to type weird stuff like this: `07 GreatVariableName PIC X(10).`
Oh yeah, and the compiler cost me $3k (great job, Micro Focus), on top of the Solaris machine I had to buy to run it. (This project happened to use `mmap()` via FFI, so it had to run on a Unix machine...)
Maybe I'm a bit bitter, but I will never touch another line of COBOL again.
>>Want to declare a variable? You're gonna have to type weird stuff like this: `07 GreatVariableName PIC X(10).`
I am experiencing this hateful pain and suffering right now. And the previous maintainers hated it too and cut and paste stuff in and out and the files are ~10k lines long and trying to track where a variable has changed, only to find that it depends on another two or more variables, and then trying to track those...
Don't do it folks.
I don't despise COBOL, it's mostly a reporting mindset language. Fitted with record file systems you can do some stuff. But if you step outside .. god help you. And if you use X-to-COBOL compilers, then even god cannot help you. Pages and pages of mangled COBOL terms. People show you printouts like war scar tissue on your first day to assert senior status :D
And if you use X-to-COBOL compilers, then even god cannot help you.
Exactly this! I started my professional carreer with COBOL, which I still don't deem to bad for what it was designed.But we were forced to use Jackson Structured Programing implemented via a pre-compiler (which conveniently ceased to work in regular intervals, since the license expired, but I digress).
You're right! God would have run screaming after one look at the resulting garbage code if we would have the impertinence to ask for help.
To answer the topic, I've tried COBOL and I hate it. It doesn't deserve respect except in a historical context.
I have to agree. Even back in 2001, I found COBOL awful. I recently moved to a new city after having worked Scala jobs for several years. All I can find here is Java work and I dread going back to the god awful word of Spring Boot+MVN+Java after having worked with a considerably better language (usually :-P).
We have better tools today, but the world is left in legacy. So few things have unit/automated tests and without a good crew of engineers, it's very difficult to move things quickly and smoothly to new systems and platforms. Just because COBOL is still in use in so many systems doesn't make it good. It just got popular at the right time, everyone embraced it, and now people who still do it are paid insane amounts of money just to maintain it.
Here's one of my old COBOL programs from University. It was written 15 years ago. I converted from CVS to Git, so the history dates are literally older than github, which is kinda funny:
https://github.com/sumdog/assignments/blob/master/csc3620-pr...
The question is, if you are not forced by a customer (such as a bank) or by the need to maintain a legacy system, why would you willingly choose COBOL?
So I don't hate COBOL, but I'm not particularly interested in trying it either.
Years later I was between gigs and interviewed with a company that wanted me to go back to COBOL. I would have, but they lo-balled the salary, like I was going to take a $20K pay cut to go back to COBOL! Ended up staying in the Java world
What I do find interesting is the attitude of some who classify COBOL as ancient (etc) and have not yet realised that there are a number of major modern languages that are the spiritual descendants of COBOL. These being ADA, Java, C#, C++ and the rest of their ilk.
They each have a place to play in the computing world, but they each are the wrong language for many purposes. Too oft I find, many programmers know one or two languages and they are the worse off for that. Different languages give different insights into specific problem domains. Every language allows some things to be done easily and other things to be done hard.
Decades ago now, as CompSci undergraduates, we were expected to learn (or at least use) every available language, from Fortran to COBOL, Algol 60 to Lisp, assembler to Simula, Snobol to Pascal, Basic to Algol 68, etc., etc., etc.
These days, undergraduates seem to be taught one or two languages and that's it. It shows up in their later inability to solve certain kinds of problems in a problem space specific way.
But programming evolved and Cobol did not evolve at the speed of the times. In 1968 (Cobol-68) it was already outdated compared to IBM PL/1 (a big monster of a language, similar but much more powerful).
And in 2017, "Try COBOL"? You mean, try a language where I need to deal with record access as if I were rolling my own DBMS system, with almost no module support, and copious need for copy/paste of code at every file (little chance of code reuse)? Where separation of concerns is almost zero?
No, i think writing in INTERCAL-72, Brainf_ck or LOLCODE would be more fun.
That gets to be a little old.
Sincerely, CHRISTOPHE (or, as I'm usually known, Christopher)
My overall impression is that it's verbose (obviously), but it's also fairly straight-forward and easy to read even if you don't know COBOL.
The most interesting part of it, for me anyway, are the 88-level declarations[0]. Basically it's like a logic test encapsulated in a variable. You can say something like "IF STAFF-MEMBER" where STAFF-MEMBER actually means that some field has a certain value. It makes it easy to understand in some respects, but also adds a layer of obfuscation.
[0] http://www.mainframestechhelp.com/tutorials/cobol/cobol-88-l...
But to be honest I'm one of these people who up till now doesn't even consider it a choice. So don't expect too much support or understanding from my side.
I don't think that's a desirable goal. COBOL is not a very good programming language: it's verbose, clunky, cumbersome and outdated. It still exists because of tons of legacy systems and (understandably) conservative businesses like banks. But if you start a new project, please DON'T consider COBOL. To be honest, even my recommendation is irrelevant: you wouldn't start a new COBOL project just as you wouldn't start a new GW Basic project -- it'd be unnatural.
Disclaimer: worked with COBOL for a bank. It wasn't pretty.
if shipping-method <> 'FX'
move normal-ship-date-yyyymmdd to expected-shipping-date
else
move nextday-ship-date-yyyymmdd to expected-shipping-date.
And not the following? if shipping-method = 'FX'
move nextday-ship-date-yyyymmdd to expected-shipping-date
else
move normal-ship-date-yyyymmdd to expected-shipping-date.
The latter seems clearer to me, and the line below (regarding the cust-type variable) suggests the syntax would be okay. It would also have the advantage of avoiding the column-length issue described in the post. What am I missing here?You want horrible? Check out RPG (Report Program Generator).
It's - Too structured. - Poor tooling. - Whitespace sensitivity. - No first-class functions.
But the article is right - if these same standards were applied to today's languages, things would be a lot louder.
It would be so cool if you could go "ADD ONE TO FIVE".
https://everything2.com/title/Changing+the+value+of+5+in+FOR...
Shipping price conforms to a "pic 99v99". So what happens if a really big shipment needs to be made for which $100.47 needs to be charged?
Having said that, situations like would come up sometimes. You really did have to decide ahead of time how big field could be. You could leave extra filler bytes in your tables (literally named "filler") but it was still a big hassle to deal with.
Are you sure about that? Sounds like rather a lot.
display "am I having fun yet?"
catch e as type Exception display e::Message
end-tryadd x to y giving z
become something like
cobol_add_giving(&a, &b, &c);
which is a lot less readable.
COBOL isn't so bad when you use a modern variant that purposefully deals with many of the complaints surrounding it.
Isn't that basically how JavaScript took over the server space via node.js?
https://github.com/azac/cobol-on-wheelchair/blob/master/READ...