Automated Refactoring of a U.S. Department of Defense Mainframe to AWS
aws.amazon.com
aws.amazon.com
Let us stop and give some caring thoughts to the people who will maintain that code base.
Certainly there are cobol to jvm bytecode options and I'm sure they see some use. One other option I haven't seen yet is cobol to an intermediate DSL in Java that legacy developers could use and code with little trouble, Jobol if you will. This doesn't exist that I know of.
You need to find that Cobol developper, that can work at your office, that will still works for at least the next 10 years, that won't ask a ton of cash because he expect to keep a similar salary as he did before, that is actually competent, etc...
In the other hands, you got 10 000 unemployed java developer that get out of school in less than 500 km around your office, every year, that will do the job nicely for a pretty cheap salary because it's their first job.
> Once modeled within the tool engine, SMEs employ an iterative process of applying rules and tuning to output the transformed code into the target Java language.
Seems like they manually construct an IR of the source system, then translate the IR to Java Code.
Automatically transpiling COBOL code to Java will lose all possibility of maintenance moving forward. So it is not considered in the first place, nor it is what they need. They need a semi-auto REWRITE of their legacy system, not just target it to Java.
even if it $100 millions, it seems a successful project , given DoD mentioned potential savings of $25 millions/year
>Customer Benefits
> For the DoD, the component has been transformed from an expensive mainframe COBOL legacy system to an affordable, modern maintainable Java-based system. All valuable existing business rules have been preserved in the modernized system, while development, test, and production environments were migrated to AWS providing flexibility at reduced cost.
>The DoD’s projected cost savings is $25 million per year, and they are now able to use popular Java programmers to maintain and enhance the critical component.
It was interesting while I was there helping to transition the Blekko stuff to IBM I brought up how this sort of modernization of their cloud offering was in their strategic interest. IBM is so big, and so complex internally, it was kind of like shouting at a crowd moving through Grand Central station in New York. People can hear you shouting and slowly start to comprehend what you're trying to get them to see. Change is hard in an organization of that size.
They already did that, they bought SoftLayer, but it's not working very well since instead of going for "a cloud offering" they are doing a "same kind of ibm offering as usual, but in the cloud !", there is a reason no one is talking about using them on sites like HN their entire target market seems to be company moving from big "enterprise" hosting to big cloud "enterprise" hosting.
And they're not buying AWS any time soon.
Here's one article from last year: AWS is now cheaper than SoftLayer https://thehftguy.com/2018/11/13/2018-cloud-pricing-aws-now-...
IBM has been around for over a hundred years. That is important because understanding how they have survived can give you an idea of the way they think about things. At several layers they tend to think in very long time frames.
I got to meet a number of the Softlayer managers when our startup was being integrated (Softlayer had people going through the same program) and I got to meet more of them when we did a project to host the Blekko crawler in a Softlayer data center on a number of machines there.
That was a good example for them of a project where the "Colocation/Hosting" model really fell down on the floor. A crawler needs (for example) a bunch of machines that have low latency access to each other as well as the Internet in order to support their storage and indexing models. AWS has 'service layer' clusters where you can subscribe to an "instance" of an elastic search cluster, which then allows you to talk to the service layer while the data center can optimize the communications layer. This is "Web 2.0" architecture which Google and Amazon both developed as part of their data center scale computers effort. The "policy" or "application" layer can be hosted on any machine because it needs only to dispatch requests of the underlying services, whether it be a storage service, an indexing service, or a network crawling service.
Back in 2014 when we were acquired this was still 'new stuff' to Softlayer and they were getting hammered both by our integration needs and the nascent Watson developer cloud for similar reasons. That was combined with billing and finance challenges (IBM is driven by the Finance group :-)) in order to understand costs, expenses, and ROI.
From what I could see from where I was sitting, it was at least a 5 to 10 year effort for them to get to where AWS was a decade ago in their thinking. But advancement from that point on could be more rapid by taking advantage of what had been tried and failed before. So many pieces of the company needed alignment, sales, finance, engineering, management, and operations. To be honest, if I was going to spend the end of my career somewhere it would have been a good project to take on.
So when I read the comment "it's not working very well" my response is "yet" and I agree that that "yet" can be a fairly arduous process within a company the size of IBM.
> how they have survived can give you an idea of the way they think about things
I don't care. The speed of innovation will outpace IBM forever, at this point. I'm not going to leverage them in my lifetime, nor wold anyone else who is competent. There are lots of bottom feeders and IBMs spectacular age is legacy entrenchment, not a magic sauce.
This can actually be said of almost every “modern” technology now and hot startup now.
I have been involved in a somwhat similar project where we had a large system written in an obsolete language. It was HUGE. And it ran fine. With millions of transaction. But it required expensive hardware to run it.
The decision was made to convert it to Java. The effort took years, and the rend result was not impressive I thought.
In a lot of ways Java was a step down in abstraction from the more DSL original langauge. Implementing new rules took more time, and running it took a lot more orchtecstration.
I am not convinced that it would have been better to hire and train a few devs in the old system and let it keep doing its thing for a few decades more.
COBOL is a very niche language that is very well specialized to banking. I don't envision applications rewritten in something else easily for any significant benefit.
But maintaining and extending the thing was only going to get more expensive. And once you've paid the price to modernize, you've bought yourselves a lot of savings for 1-2 decades, maybe more.
We kinda see something similar today with Tesla - and we'll probably see more of it in the future.
And they ripped out the COBOL Data management system and replaced it with a RDBMS. This alone probably is worth it since you can hook it up to business intelligence tools now to create reports and dashboards.
My concerns mainly lie with choosing such a proprietary solution over one using more open-source standards. Both in the language selected, along with the backend (Oracle and AWS) being used.
Will there or could there be problems in the future, should the need arise to migrate off one or more of those proprietary platforms?
What will or could happen if there is a security issue that affects or targets AWS, Java, or this system in particular - are we "stuck" again with a potentially "broken" system that can't be easily moved because of vendor lock-in of sorts?
Should Oracle make further moves in the direction of "closing" Java - will programmers continue to learn the language and support it, or will they move to other, more open, solutions?
What happens if or when AWS or Amazon ceases to exist?
I just wonder if we haven't traded a largely unwieldy but mostly known issue of a mainframe and COBOL, for a seemingly known (but questionable) issue of AWS, Java, and Oracle.
Will we just revisit this whole problem again in 50 years?
If so, maybe then we'll make a better decision to use more open and auditable solutions (but I'm not going to hold my breath to that)...
Fair point given how things usually go in the private sector. In the public sector, especially DoD, the rules of the game are changed.
What happens if or when AWS or Amazon ceases to exist?
Well guess what, they can't. It's now a national security issue so the necessary services will continue to be provided by a KTLO crew or DoD would facilitate a transaction where another company could absorb the needed technology and be contracted to support it. DoD is the biggest company on earth. FY 2019 budget is $686B.
DoD has a pattern of keeping mission critical end-of-life systems running and supported.
I'm currently advising the US Navy on moving some COBOL systems to Java. Those systems have been lovingly supported by lifelong Navy civilian employees for 40-50 years. Those people are now in retirement age and so they're "replatforming" to Java.
There really is no corollary in the private sector. The default for government systems has been the people who built it stick around and maintain it until they retire. For DoD the biggest risk isn't old technologies and changes in the private sector, it's brain drain when people retire.
You might ask yourself, what 20-something wants to work for the US Navy, as a civilian, and work on a Java implementation (ported from COBOL) for the rest of their career? Well, I'd like to introduce you to them. These guys could work anywhere but they like the Navy. They like that their work matters, it's important, and it's stable. The pay is below market but it isn't terrible. Keep in mind DoD isn't all in DC, many offices are in inexpensive places to live.
Frankly I'd love to work on systems with the meaningful impact that theirs have (non-combat in this case). Read about 18F and what it was like for the people who joined and how they stuck with it anyway.
This underlines how so much of the work and contracts at the DoD don't really revolve around defending the US than being another way to hand out tax payer dollars to already well off corporations.
https://federalnewsnetwork.com/tom-temin-commentary/2018/01/...
edit: Added link to patent application.
> [Wang] says he applied to become a GS-15 or Senior Executive Service member so he could see through the assembler-to-Java project. But his approval didn’t come through until a week before his employment authority expired.
They lost a guy responsible for software that could save taxpayers how much -- tens of millions of dollars? More? All because they couldn't pay a guy a GS-15 salary, which translates to $100k-140k.
Man, people are really dumb about paying programmers well. They'd rather see a project go off the rails than pay someone (still slightly below) market rates.
Also:
> In many ways, assembler is still excellent for this application. Milholland said of the code, “The assembler is well written. It’s incredibly efficient and effective.” But a shrinking number of people understand it. And it’s not optimized for the online, transaction mode to which the IRS needs to keep moving. Java, relatively inefficient as it may be, is the current standard and has legions of people who know it.
That's like arguing mixing concrete by hand is "more efficient" because doing so allows you to be more frugal with concrete.
No, that's more like saying mixing concert by hand starts to become attractive because there's only one guy on Earth who knows how to operate a cement mixer. The alternative is crap, but it's easy to find people to do it.
This is the reason that makes me believe Go is going to be the new Java: a language with good enough performance, lots of safeguards preventing sloppiness to break, and a very very flat learning curve.
Recompiling the COBOL codebase on Linux and accessing and loading it as shared objects (via JNI for JVM usage)
Running Hercules ( https://en.wikipedia.org/wiki/Hercules_%28emulator%29 ), an IBM mainframe emulator
The first one has lots of merits, as it allows for a progressive replacement of COBOL functionalities
The second one immediately lowers the TCO for obvious licensing reasons but does not plan for the future, particularly as COBOL developers retire and thus the resources are scarce.
(Anecdotally the customer chose both: recompilation & progressive replacement by Java code for core assets and emulation for "dead", non core, assets dev & staging environments)
How true is that though? They'd still have software licensing, they still need an OS to run on Hercules after all.
Linux sysadmins were cheaper than their mainframe colleagues.
Hw was orders of magnitude cheaper.
Floor occupancy was lower.
We did not got as far as watts consumed per functional task performed but I doubt it would have tilted the balance?
You're obviously right for licensing inside the emulator! However, the costs are(were?) in hardware & floor occupancy (as in "decomission a whole room"-cheaper)
The COBOL doesn't have automated tests. When you port the COBOL to Java (whether automated or manually) you aren't asking "where are the original requirements?". Contrast this with a "re-write". When you do a re-write you have to go back to the customer and whatever documentation you can find and figure out and re-document what the system should do. This can be far removed from what the system did in its original design docs, and can differ still from what people believe.
This step of codifying tacit knowledge of the behaviour is important in the evolution of systems.
You should be asking where is the current documentation of the system specs, and in some cases you might even get something reasonably current and complete when you ask that, depending on how to use the change management process has been.
> When you do a re-write you have to go back to the customer and whatever original documentation you have lying around and you have to figure out and document what the system does.
Both what it does and (and this is where the customer is involved) what it is currently supposed to do.
Another possible target for this is the IPPS/OPPS Medicaid/Medicare code calculator system - which was written in COBOL for IBM 370 by mostly a single developer. Sadly, the developer passed away about 10 years ago, and the system has yet to be updated. The code was open sourced in a bid to gain assistance for updating from the community. Automation such as this would help significantly.
The code and data files are all here, enjoy :)
https://www.cms.gov/Medicare/Medicare-Fee-for-Service-Paymen...
They spent 12 months making it less-COBOLy / autogenerated feeling
There are many many problems with these types of projects. One thing that often happens is leaving the new application in the hands of the legacy team to manage. They often lack even rudimentary skills in managing java runtimes. I saw one team struggle with performance issues that were easily solvable with a connection pool configuration change.
Maybe the worst case with these types of systems is enhancements going forward. You can guarantee there will be no more money spent on modernization, and any changes will be made to the Frankenstein java/cobol source. Which will continue to accrue technical debt and an even greater rate.
The main project I saw that did this was eventually abandoned, and luckily the cobol application was still running.
It's still a win if they can go from procedural COBOL to procedural idiomatic Java.
I saw one team struggle with performance issues that were easily solvable with a connection pool configuration change.
Seen it happen before. Someone will step in and charge those hourly consulting fees.
You can guarantee there will be no more money spent on modernization, and any changes will be made to the Frankenstein java/cobol source. Which will continue to accrue technical debt and an even greater rate.
Not if the project includes a phase where they start modifying things to increase the degree of idiomatic Java code.
Goto's and "fall through PERFORMS" make idiomatic java impossible.
No. The AST part is wrong. Compilers do decompose code into basic blocks. Even the most spaghettified code decomposes into basic blocks. Any language with tail call optimization can handle the basic blocks rewritten as functions.
That won't get you to "clean" however.
It'd be interesting to compare this approach to modernization approaches in the government. For example the veterans administration is moving away from their ancient but popular and decently well regarded electronic health record system Vista (https://en.wikipedia.org/wiki/VistA) to cerner.
https://www.hitechanswers.net/va-establishes-office-of-ehr-m...
The other hidden cost is that these ancient software systems are extremely efficient in terms of cpu / memory utilization compared to modern equivalents so the hardware requirements go up considerably which impacts power / energy efficiency and direct deployment costs for replacement systems. Not saying modernization is not a win in the long term but there are direct and indirect costs with these big transitions.
Great use to match up with the dirty Java logo. I’m sure they had a lot of fun designing that, and this dirty/clean terminology will probably stick in other source to source transformations.
Next step is to fix the bugs, and get all the tests to run.
The last step, to get to "clean" involves going through the code base, looking for project-specific idioms, then hand-tuning the syntax matching engine with rules that output idiomatic target language code. This might also involve refactorings to make the scaffolding libraries more idiomatic as well.
Since one is coding on the transformation engine, instead of directly on the application itself, production changes and maintenance can proceed, and no code freeze is needed.
(this rarely works out well for taxpayers in the long term..)
I did find it interesting that the original system, and delivered overhaul, both ran Oracle databases. Overall transactional volume is low (500k/day). Perhaps the DoD has significant investments in stored procedures? Or perhaps Aurora is not ready for prime time in AWS Gov Cloud? (It’s no secret that AWS really struggles with keeping feature parity between Gov and commercial.)
Maybe the business purpose was to waste time and bill clients?
They mentioned numerous times that COBOL work was harder to staff and maintain.
Seems like a good reason to move to me.
As one example, when I was a systems engineer, I got a call from a team that was trying to run a business application from z/OS in emulation on an x86 server. It was running incredibly slowly, even on a fairly beefy server. I looked in its data directory, and there were hundreds of thousands of individual files in a single directory. Not a tree. One directory. That might work fine on a mainframe, but just doing a dir or ls (depending on the host OS) took something like ten minutes.
Similarly, I remember some folks being dumbfounded at the idea that their batch processes needed error-handling when they were doing things that involved talking over the network, because when everything was running within that one giant IBM system, failures were pretty rare.
Those both sound like they should be easy to tack on when running in an emulator, but they're just the tip of the iceberg. When you have code that's 30-50 years old, written for a vastly different platform, there are going to be a lot of gotchas like that.
> "Difficulty finding qualified, affordable COBOL programmers."
Continuing to emulate the COBOL environment doesn't quite solve the problem that finding COBOL developers is becoming harder and harder.
Anyone else experience some cognitive dissonance in "the world's most customer centric company" also someday powering the facial recognition systems used by customs, border patrol, police, and military systems?
It seems ripped straight of a cyberpunk novel to me.
disc: I'm Amazon employee, not working in any of the facial recognition. I'm also an immigrant, ~5 years in USA
That's just a lone hacker doing shit outside of work... And my solution can go through directories of images, realtime video stream, videos, and ipcams.
So yeah the cat is out of the bag. The bag was shredded, and the cat has had 3 litters of kittens.
Any video system could be retrofitted as a facial recognition system... Or even process videos after the fact.
I'm a bit surprised DoD encourages the consolidation, though maybe once the software is cloud enabled it will be more portable to small sites in the future.
AWS has 13 US-based availability zones, each of which has, as I understand, multiple DCs.
GCP has a similar architecture with 16 US-based zones and 3 more standing up imminently.
Azure has a similar architecture, but separate AZs within regions are newer for Azure, and I can't find detailed information as readily, but they look to have 8 US regions,at least four of which are multi-AZ, so at least 12 zones, each with one or more DCs.
And that's 41 zones, many of which have multiple DCs, with just three big public cloud providers; I don't think 10 DCs gets you most of the computing power.
GovCloud East is also up, but the GovCloud West region's 3 AZs plus GovCloud East aren't all of US (or even “US government on public cloud”) compute (AWS also has Secret Region, Azure has two government non-secret regions and two secret regions; GCP doesn't have a government region but federal government agencies use regular commercial GCP regions—that’s true of AWS and Azure, too, the government offerings are approved for high-impact use whereas the general offerings have a lower-impact approval; Google is seeking to also get the High impact approval for their commercial offering, rather than having a distinct region IIRC.)
https://www.google.com/about/datacenters/inside/locations/in...
Quoting from the docs [1]:
Each AZ can be multiple data centers (typically 3), and at full scale can be hundreds of thousands of servers. They are fully isolated partitions of the AWS Global Infrastructure. With their own power infrastructure, the AZs are physically separated by a meaningful distance, many kilometers, from any other AZ, although all are within 100 km (60 miles of each other).
All AZs are interconnected with high-bandwidth, low-latency networking, over fully redundant, dedicated metro fiber providing high-throughput, low-latency networking between AZs. The network performance is sufficient to accomplish synchronous replication between AZs.
1 - https://aws.amazon.com/about-aws/global-infrastructure/regio...