Edit: and when do you upgrade? If they upgraded in ‘85, it would be in C++. ‘95, and it would be Java 1.0. ‘05 and it would have been VB6. None of those would have been substantially better save that they would have been easier to hire maintenance programmers for.
Maybe write the tests first so it can be validated easier?
With so many legacy systems that are often in very important places, I wonder whether it wouldn't be smarter to spend money on systems converting a modern language to cobol, e.g. python2cobol. Is that impossible?
The closest that modern design patterns come to these systems, is using them as the nightmarish example that justifies why modern practices exist.
It's rarely actually a technological problem. It's scale, scope, documentation, budget and motivation. You have to take Mount Everest, and carve it into 2 million separate boulders. Document every single one of them. Paint some of them. Replace some of them with stronger materials. If any of them move, you failed. If any snow is disturbed, you failed. If the climbers even notice this is happening, or has happened, you've failed. And on top of this herculean feat, the person paying for it needs to understand that despite the insane cost of this endeavour, he's probably not going to see a single benefit - but his successor in 10-20 years will. But if you fail, he's going to feel that hard and fast.
Cobol was a modern language 50 years ago. Python will be an ancient language 50 years from now. What do we do? Keep rewriting everything every 50 years?
And those ancient Cobol codebases actually have a big advantage: they've been maintained for so long that all the major bugs have been virtually eliminated. Creating a new system from scratch means another 10 to 20 years of maintance until the new system reaches the stability of the old one.
This is no joke. Given that most of the Cobol running today is running on mainframes at banks and card networks and the like, "a new bug" may translate to a few hundred thousand dollars of losses.
There's no point in building a new house of cards every few decades.
Uh… sure? What would you consider an acceptable minimum amount of time to be after which rewriting a large but fairly critical codebase to modern standards becomes acceptable?
> And those ancient Cobol codebases actually have a big advantage: they've been maintained for so long that all the major bugs have been virtually eliminated. Creating a new system from scratch means another 10 to 20 years of maintance until the new system reaches the stability of the old one.
That's not quite fair. Having a known good code base while porting means that you can just rewrite the existing algorithms in the new language and then run some automated tests to make sure you get the same output from the same input across both systems. Unless you're doing a black-box rewrite for some reason, you're not really throwing away all fo the maintenance done on the existing system.
Re-writing an existing algorithm is one thing, but even that is likely to be a big source of new bugs, given that Cobol is actually a quite low-level language and much code will rely on its specific view of a mainframe's architecture.
The bigger problem is that any implementation of complex business logic is going to depend very heavily on the facilities provided by whatever language it's originally implemented in (Cobol, in this case, obviously). A direct translation to a new language is likely to be completely impossible. And the bugs will grow in all the semantic gaps between the old language, and the new.
And that's before considering that, for Cobol in particular, the Cobol code itself is only half the story. Cobol programs run as batch jobs controlled by JCL ("Job Control Language") which often means that crucial aspects of businees logic are spread over multiple files in _two_ languages. And the JCL part is a mess. I didn't mind Cobol when I was working with it, I even came to like it a bit actually. JCL is really, really awful.
But, aesthetics aside, where does all the JCL-encoded logic go? Is that translated to the new language, also? That's going to be really hard given that JCL is operating-system specific. Is it going to be translated in scripts in a new shell language? The difference between concepts on JCL and, say, bash, or powershel, is going to be impossible to bridge without making drastic changes- and cultivating new bugs.
In general, translation of a large codebase between two very different languages is going to cause lots and lots of new bugs. So, if you rewrite everything every 50 years, in 200 years you'll spend a total of 40-80 years fixing bugs. If you write it once and let it be, you'll spend at most 20. I don't see a good reason to do it.
And what's wrong with an "antiquated language" anyway? I mean is it just aesthetics we're talking about here? Is it the lack of programmers that's the problem? The latter is sure to make translation even harder and more bug-prone. What is the real reason to change a working codebase every n years?
Yes, or sooner.
And you keep the 'institutional knowledge' externalized in documents, and you use testing tools and use virtualized systems and what not (virtual systems were available to consumers even 20 years ago - this is not that new).
"Upgrade" or "rewrite" every 10-15 years. This should just be a cost of maintenance. I'm at the point where I've had PHP code running on systems for 15+ years (had a call from someone in 2017 about software started in 2002 and last touched in 2004). There's a 'on the public internet' distinction with web apps vs internal bank systems, for example, agreed, but it doesn't remove the need for upgrading old systems. Doing it on your own schedule, on your own terms, vs having to deal with systems in crisis, is where the benefit is.
"what not" has been running in production since 1972: https://www.ibm.com/it-infrastructure/z/zvm
Somehow I think C or C++ wouldn't run into the same problem while I can see this happen to python.
http://www.csis.ul.ie/cobol/examples/Conditn/Conditions.htm
identification division.
program-id. letters.
data division.
working-storage section.
01 Char PIC X.
88 Vowel VALUE "a", "e", "i", "o", "u".
88 Consonant VALUE "b", "c", "d", "f", "g", "h"
"j" THRU "n", "p" THRU "t", "v" THRU "z".
88 Digit VALUE "0" THRU "9".
88 ValidCharacter VALUE "a" THRU "z", "0" THRU "9".
procedure division.
begin.
display "Enter lower-case character or digit. No data ends.".
accept Char.
perform until not ValidCharacter
evaluate true
when Vowel display "The letter " Char " is a vowel."
when Consonant display "The letter " Char " is a consonant."
when Digit display Char " is a digit."
when other display "problems found"
end-evaluate
accept Char
end-perform
stop run.
(I removed some chaff which GnuCOBOL doesn't need and fixed an apparent bug.)So... where is it actually testing what kind of character you input? Where is the code for that? You input a specification for what a Vowel is, for example, and you write explicit code for what to do when a Vowel is input, but where is the code which goes through the specification and decides, yep, that's a Vowel?
COBOL is kind of an odd language. It's verbose in some respects and quite concise in others. Rewriting COBOL into something else would take actual human effort if you wanted the "something else" to look like code a human wrote, as opposed to the intermediate pass of an optimizing compiler, which is what the C GnuCOBOL can output looks like. Re-writing the COBOL might be the best move in some cases, or replacing it with entirely new code, but it isn't something you'd be able to do "for free" in any sense, especially with regards to time.
Cobol is really not such a bad language. It's just got a lot of ...ceremony. All those forced divisions and sections. But that's a feature: in the olden days, structured programming was a big thing. And an experienced Cobol programmer can take a quick look at a big Cobol file and find where everything is in a blink.
It is, but the "magic" is that the pattern-matching is part of the variable declaration, so you can reuse those patterns wherever you can use the variable.
sub letters
dim char as string
do
char = inputbox("Enter lower-case character or digit. No data ends.")
select case char
case "a", "e", "i", "o", "u"
msgbox "The letter " & char & " is a vowel."
case "b", "c", "d", "f", "g", "h", _
"j" to "n", "p" to "t", "v" to "z"
msgbox "The letter " & char & " is a consonant."
case "0" to "9"
msgbox char & " is a digit."
case else
exit do
end select
loop
end sub
(This example is actually LibreOffice Basic, as I'm not on a Windows machine right now, but it should be the same in VBA and VB6.)More to the point, it allows you to keep the patterns near the variable declaration, so you can reuse them.
The pattern matching is part of how the data is declared.
Aside, RE testing, a wild guess: could the "accept char" be tested with emulated keyboard inputs?
Govt IT systems and procurement are such a mess, putting it all out there for people to review and complain about is the only way it'll ever get better.