For those of us with only exposure to PCs and commodity servers, I'd be interested to know who has encountered one in their day to day, and how that experience differed from the norm.
For those of us with only exposure to PCs and commodity servers, I'd be interested to know who has encountered one in their day to day, and how that experience differed from the norm.
A system that runs a bank or an insurance company and wants good availability can be built in one of two ways: you either spend money on software that deals with the hardware being unreliable and save money on software (pioneered by Google) or you spend money on hardware that promises to be highly reliable and save on software.
No new player believes mainframe (extremely expensive hardware) is cost effective, they all use commodity hardware.
A bank that needs to run binaries from the 70s for which they don't have the source code can keep paying IBM and not investing in reverse engineering the binary and implementing it in Java.
A bank that has a billion lines of code in Cobol can compile it to run on jvm on commodity hardware and run it in parallel with the mainframe for a year to validate and then switch over to the new system and stop overpaying for hardware, but that sounds risky, so they keep paying a million dollars for a system with the same performance as a 50 thousand dollar server.
Seems like the first "save money on software" shouldn't be there.
That approach is, in scientific computing, at least as old as Beowulf (1994, four years before Google was founded), the prototypical system from which we get the term “beowulf cluster”.
Edit accidentally included runaway thought process...
It's OK to do most things with chunk of normal servers, but when you need to handle very large number of transactions with fast commits, and stuff like eventual consistency is not allowed, it becomes expensive to handle them no matter what.
One of the other big Australian banks (Westpac) has an "exit IBM" project as well, but it isn't complete yet.
There is no information anywhere saying that they moved their critical banking systems and transaction processing.
Unisys always had a smaller mainframe share than IBM, and they stopped making new chips sometime in the 2000s.
For at least ten years the fastest Unisys mainframes have been very big (proprietary) x86 machines running Windows with a Unisys mainframe emulator.
For some time I defined a "serious computer" as something that didn't have ports for keyboard and monitor.
Having said that, the pre-x86 A series were pretty cool. And ran the most user hostile OS ever created, to the point its very name was used for Tron's villain.
Monzo? They're getting bigger now, circa 3M customers. More than First Direct, Starling, Metro Bank.
I personally worked on a project there that involved building new virtualization platforms for not only x86/x86_64 Intel, but also AIX (Power) and Solaris (SPARC).
I was told they did an assessment a year or two ago to price out what it would take to move the business completely off mainframe and onto an x86 stack. It was in the hundreds of millions of dollars to do so because so much other software has been built to interact with and rely on the mainframe over the years that switching off it would be a multi-year effort across every department in the company. So of course that ROI calculation was pretty damn easy and the mainframe isn't going anywhere.
So just... have the new server expose itself over the TN3270/TN5250 protocols, such that these interoperating systems see what they expect to see? (You could even build the new system as a regular REST API or something at its core, and then build the TN3270/TN5250 exposure as a separate gateway service on top, such that it'd be easy to shut it off later if everything finally moves off it one day.)
You don't necessarily need to move the $new project off the legacy hardware, just write it in such a way that you can do easily later.
I'm eliding many details here, but the principle stands.
- Spend $300,000,000 over 10 years moving to a new system, and hope that by the time you've done that it's not obsolete.
- Spend $1,000,000 a year on a system that still works.
You could run the system for 300 years for what it would cost to replace the system.
In spite of the collective wisdom on HN, there aren't a lot of companies in the world with hundreds of millions of dollars sitting around doing nothing. Even ones that work on mainframes.
- Spend $1.5 million maintaining the current system, but as bits get updated keep in mind the system that you'd like to have in 10 15 years time.
You're going to have to replace the system within 300 years anyway, you aren't saving that money, every feature you add that is reliant on the old system is literally technical debt, because eventually you'll have to rewrite it for the new system.
I'm not even saying you need to have a new system in mind, just keep in mind that you will be moving to a new system, so code appropriately.
And it will still be able to run all your programs that were written and tested since the late 20th century, by the kind of organic entity we used to call "human".
Slightly more serious retort: I'm not aware of any brands from 300 years ago, I'd be surprised if any of IBMs customers survive, let alone enough to keep IBM as a going concern.
Btw when you start folding the space time mesh, terahertz figures just become marketing numbers, what you really want to know is how many parsecs it can do the Kessel run in.
You might be, you just don't realize that they're hundreds of years old.
For example, the insurance company Lloyds of London is fairly well-known around the globe. It's 333 years old.
If anything, this is a sign that mainframe users have way too much profit for what should probably be commodity software.
Remember that these are large, public businesses. Explaining to shareholders that profits are going to take a noticeable hit for years because of IT investment that isn't strictly necessary is effectively a non-starter.
I wasn't thinking that. I was thinking more like when you add a new feature, fit it into an API that is portable. Or add a translation layer so the feature can be written how you would like system to be in the future, but it works on your hardware today.
2nd I'd say theres a half life of best practise. Over the 10 20 year tome frame, id expect some of what you were doing is going to become outdated yes, on the same way some of your knowledge over your career will become outdated, or the phone I your pocket will, you wouldn't use that as an argument against education, or buying that phone.
Btw, if its as costly to deal with the interface as with the mainframe directly, that's a win. The interface can move off the mainframe to somewhere cheaper/better.
It doesn't become wrong, it just evolves into another dead-end. The 3270 is an interface, too, but the reason everybody suggests replacing it or augmenting it is precisely because the spartan tooling and mindshare makes it expensive. As XML recedes into history it is likewise becoming more expensive as an interface.
I chose XML vs JSON because I figured it was a transition everybody was somewhat familiar with. And younger programmers have an almost visceral dislike of XML, which I thought might help get the point across--that an interface someone once thought (and probably still thinks) would help ease future interoperability becomes a reason or excuse for future programmers to avoid that integration.
> Btw, if its as costly to deal with the interface as with the mainframe directly, that's a win.
I think the problem is that you don't really know if it's as costly. The error bars on that sort of risk assessment are huge because our industry sucks at accurately predicting migration costs. And it sucks because complex software systems are intrinsically unique. Commercial solutions that claim to be able to capture and control all those dimensions of complexity tend to be sold by vendors with names like IBM and Oracle. Such vendors also pioneered interfaces like SQL, which is both a soaring achievement in terms of capturing complexity behind a beautiful interface while also falling epically short of what's needed to actually reduce long-term integration costs.
But alas, language has changed and these kids won't get off my lawn.
‘Whatever is fitted in any sort to excite the ideas of pain, and danger, that is to say, whatever is in any sort terrible, or is conversant about terrible objects, or operates in a manner analogous to terror, is a source of the sublime; that is, it is productive of the strongest emotion which the mind is capable of feeling.’
What is a problem is that the last time a lot of code was touched may be 5-10 years ago. That code may have started being written 20-40 years ago. It may well have been maintained by people who think "if it was hard to write, it should be hard to read" or "documentation is for the weak". It definitely will have been written when the cost of a gig of memory and storage was many orders of magnitude higher than today (indeed, last time I priced memory for a mainframe, a Z10, it ran to $10,000 a gig); hence terseness in everything from table and column names through stored data and everything else was prized. Dropping from COBOL into assembler is not uncommon for critical path performance.
Making any changes will be a week of coding and three months of working out the what and why of the code, because the last person who worked on it retired a couple of years ago.
It has worked for the past 30 years and hasn't been touched since 20.
Who are the vendors that provide a migration path from mainframe to a set of commodity hardware running an emulator, and subsequently provide continual maintenance and support?
Global 100 company gets IT from other global 100 company
vs
Global 100 company gets IT from small mainframe support shop
there's a question of shareholder liability, being able to adequately sue them for M's of $, expectation they will be around in 20 years, etc.
A single computer with 4 sockets of xeons will outperform a mainframe but will have more downtime.
The mainframe has best possible single thread performance and as much cache as possible, redundancy and parts can be replaced while it's online, but not that many cores.
The cost is very high - when I looked, 1 million per year is the baby version with only 1 CPU enabled and no license to run the cryptographic accelerator and limits on software, etc.
Commodity hardware you buy and use for 5 years, so the amount of good hardware you can buy for the price of owning a mainframe for 5 years is a lot.
For the money, you can buy a lot more CPU, ram, network, storage, etc and hire Kyle Kingsbury to audit your distributed database.
At one point, IBM was selling base-level mainframes for $75,000 (see https://arstechnica.com/information-technology/2013/07/ibm-u...)
It is true though that a 'realistic' configuration is likely to cost north of $1 million, and that none of these numbers include the price of the software.
Another is that the software and hardware designed to turn a baby system into anything useful is astonishingly expensive to people not used to dealing with this end of the market. Enjoy finding that your hypervisor is licensed at 5 figures per core, and that your cores are six figures a pop.
Even if the mainframe never goes down, the entire site will go down (eventually the rack power supply, HVAC, fiber, natural disaster, backhoe, etc. will get you even if your CPUs and RAM are redundant and replaced before they fail), and then either your entire business stops or at least processing for that region stops, or your system is resilient to site failure because you built a distributed system anyway.
If you could rewrite your software to be distributed and handle a node/site going down, you could run a single site on 5 servers that together outperform the mainframe (by a lot) and can be serviced on a whole server basis (though of course, expensive x86 servers also have reliability features), or use really cheap hardware without even redundant power supplies, but have enough of them to not care.
The modern solutions are better than the mainframe, and the only reason to use them is management risk aversion and unwillingness to learn new things.
100 miles will add 5ms (round trip) to your disk flush on commit. So a system like this has the sequential and random IO latencies of a RAID of SSDs but the flush (database commit) times of a 15K RPM spinning rust disk. People lived with mechanical disks, it's ok.
Sync disk replication (in one direction) over a fiber line is not an exclusive feature. Having both sides be active, instead of active and hot standby requires some smarts from the software, but modern distributed databases do that, and if you're careful you can get far with batch sync jobs.
For read-only batch computations you can always add some extra redundancy and partition the problem. So, I don't think it is likely that a mainframe would be useful here.
But we have literally thousands of internally developed applications. We can move thousands of apps the cloud and still have a need to keep thousands on virtual/physical machines. My own apps are stuck on commodity physical hardware for at least the next few years.
The type of applications that can historically been run on mainframes is not really moving to AWS/cloud. Most of what's going to the cloud is what I would consider to be "supporting" applications, not core applications.
My own experience, that of others may differ.
Programmers are taught that database transactions exist so that when you move money from one account to another and crash in the middle, no money is ever created or destroyed. Well, cat picture websites might do that, but banks don't. They reconcile logs at end of day.
> you either spend money on software that deals with the hardware being unreliable...or you spend money on hardware that promises to be highly reliable and save on software.
Had to perform regular backups on the system as part of my internship.
Nowadays kids are all up with WebAssembly and WASI, well that is just how OS/400 works during the last 30+ years.
Originally designed in a mix of PL/S and Assembly, everything else (RPG, Cobol, PL/I, C, C++) compiled into ILE (Intermediate Language Environment) and cross language calls are relatively easy to do.
ILE applications are AOT compiled either at installation time, or anytime some critical hardware has changed or the application themselves have been updated.
Nowadays there is also Metal C (real native not ILE), Java via IBM's own JVM (which in early versions converted JVM bytecodes into ILE ones), and the C and C++ compilers are also to target actual native code besides ILE.
The database backed filesystem took some time getting used to, coming from Amiga/MS-DOS/Windows 3.x/Xenix experience, and the command line felt more cryptic those OSes, given the use of special characters as part of the name.
The company I did my internship was using them for their accounting, everything else were MS-DOS/Windows computers connected via Novel Netware, zero UNIX flavour in sight.
IBM z and Unisys ClearPath are two other mainframe models that also follow similar bytecode based deployment formats.
So in a sense you can say that on every Android phone, watchOS or Windows PC lives a little mainframe.
Ie. it shares way more with regular server boxes than mainframes.
The main three benefits of the mainframes (as a non-mainframer):
- crazy amounts of caches
- crazy amounts of pcie-slots (and sufficient internal io and processing power to make it balanced)
- production environment mentality for everyone involved, incl. extremely engineered (and redundant) hardware
I think new customers are likely to run Linux on such a box in the future... The main cost-problem for mainframe customers is software cost (esp. z/OS, cobol etc..).
I have heard stories (from actual mainframers) about the sql-performance and key-value performance (read like mongodb) on these boxes that are eye-watering..
And when it comes to reaching things outside of the mainframe itself, generally anything you can do to make accessing the data wider is better... Why use a single 16GB/sec HBA when you can interleave 8 or more FCAL paths to the same disk. Why use FCAL when you can use Infniband, etc etc.. They have a LOT of PCIe lanes available, and enough chips to keep most of them busy most of the time.
TL;DR if it's expensive and offers more paths between things, it's probably an option for a mainframe.
Does that mean great or awful?
I would love to see an article that breaks down clustering v. supercomputers v. mainframes in tangible use-cases. It is all a bit opaque to me where one officially starts and the other one(s) begin. As well as where (and to what degree) the value prop is of one versus the others.
I've heard folks say mainframes are just legacy stuff carried forward - I've heard others say they have specific use-cases that are not well solved by other current technologies (usually a cost of refactoring). That said, are there any current use-cases, where if I were to write code from ground zero, that are best solved on mainframe, hands down? That is my root question.
There are a lot of systems that are simpler (cheaper and less risky) to scale up than re-engineer to be distributed.
For example at the end of each day a bank has to generate the final balances of the accounts that it has with other banks and then it has to check for irregularities and then finally set those balances just right (e.g. enough to cover payments made by its customers but without leaving there millions) => if it doesn't then it might not be able to use the money that is "parked" there by mistake, or it might have a huge position with a risky bank, or if the balance is too low other banks might decide not to perform the final credit into the target accounts involved in the customers' payments (which would then generate a lot of complaints + lower the bank's reputation with its long-term consequences), etc... . (the opposite happens as well - https://www.investopedia.com/ask/answers/051815/what-differe... )
Nowadays even accounting (coupled with risk management) has become time-critical because of the many regulations - a 1-day delay filing the numbers with regulators and/or central banks can result in huge fines and loss of reputation (if the news about the problem becomes public it could happen that the news services try to highlight it increasing the focus on that negative publicity). In our case the detailed/low-level data (e.g. client X bought N shares of some company, client Y retrieved N$ from the ATM, etc) is all processed by the mainframe, which works fine 99.99999% of the time => then anything that comes afterwards which involves analysis/aggregation/reporting/etc of that data happens in other non-mainframe apps (they're all very different - from tiny to huge apps with distributed databases) and in general any such app has usually some kind of problem that results in a delay at least once every quarter due to the SW itself or even the HW => it's usually not a big problem (usually some hours are lost but there is a buffer) but if the central SW (running on the mainframe) would have the same problem then there would be a chain-reaction on all the other apps that need that data (and then their specific problems would be on top of that, and then there would be a hotspot of needed 100% CPU/RAM/network resources as they would all run at the same time as fast as possible which could cause further delays, and so on).
https://medium.com/@bellmar/is-cobol-holding-you-hostage-wit...
COBOL was in effect designed for business use so fixed point arithmetic has a lot of support in the language, libraries, and tooling.
Once you do that, it becomes clear Java is not exactly a competitor.
Which is neither here nor there. COBOL has end to end support for financial/commercial calculations, and a whole lot more.
"When you are a major financial institution processing millions of transactions per second requiring decimal precision, it could actually be cheaper to train engineers in COBOL than pay extra in resources and performance to migrate to a more popular language. After all, popularity shifts over time."
But because is not in-built any advantage of use it get lost in the noise.
That's one of the best typos I've ever seen.
(Although I’m sure the precise problem set could be replicated with commodity hardware too).
Only pretty recently. You would be surprised at the number of crazy smart people that failed to make a commercially viable TPF replacement.
Including ITA, whom Google bought for ~$700M. Lots of top talent, lots of funding. They did build a reservation system, but nobody of note would use it.
Amadeus only got rid of their TPF mainframes in the last year or so. As far as I know, Sabre hasn't finished.
I don't know what progress VISA has made (another heavy TPF user). They are still hiring TPF programmers: https://usa.visa.com/careers/job-details.jobid.7439996782301...
People moving away usually tried to extract the business logic such that the TPF layer was mostly a distributed NoSQL store. That bought a lot of time, but it seems like the skill shortage is now hitting that layer.
You can certainly have a warm second mainframe too, and I think most large banks are run that way, but it wouldn't be very fashionable to add that to your airline now, if it wasn't already there.
It's a lot easier to find good engineers to write and maintain software for Linux on x86-64 than it is to find mainframe engineers, which makes things simpler, arguably.
How long is Linux on x86-64 going to stay a thing? What are you going to do about staging Linux updates on business-critical facilities?
Perhaps not so simple after all.
Which suggests an interesting contrarian career path - train on COBOL + mainframe you'll never lack well-paid work. You'll also avoid the usual fad tracking.
Citation needed, especially on that implication that x86-64 lacks reliability.
I’ve run reliable services on commodity hardware that literally processed over 5 billion transactions per day.
> How long is Linux on x86-64 going to stay a thing?
If something better comes along, why wouldn’t you want to move to it? But “better” in this context would imply popular support and plentiful access to developers, so you would have years of warning. It wouldn’t come as a surprise. As of now, it has been going for at least a decade, and shows no signs of stopping.
> What are you going to do about staging Linux updates on business-critical facilities?
This is not some huge, unsolved problem. It has been solved many times, in my opinion, so just learn how others do it and follow suit. I’m not here to teach sysadmin “tips and tricks.”
I would be completely clueless on how to stage updates for a mainframe. Turn it off and back on? And I’m betting all of the good training material there is locked behind expensive paywalls.
> Perhaps not so simple after all.
Disagree, especially since you can hire people to solve these problems for a lot less money than you can get just the hardware for a mainframe, let alone hire the extremely rare (read: expensive) personnel needed to maintain and develop applications for mainframe.
If IBM would focus on bringing the entry-level cost of mainframe down, they would probably be able to get more adoption, and more people would be able and motivated to learn their systems.
Basically, the hardware and software on a mainframe does it for you. Versus having to implement redundancy and resilience into your app. You can pull a CPU on a running mainframe, and it keeps chugging. Batch failures, app failures, etc, have a very well defined ecosystem for recovery that's consistent across apps.
As you imply, though, provided you pick the right software, there's not much difference in reliability these days. The variety of software choices is what kills reliability on x86/Linux. Too little established experience because there's too many choices.
Compared to what? A cluster of x86 computers where all the security, cryptography, observability, reliability, availability has to be written by the owner on top of something like Kubernetes? And that can achieve the kind of throughput a mainframe has?
I'd say a mainframe is much simpler than that. It's already done. All you need to do is to sign the check and read the (hundreds of) manuals.
if it's already in place, it's not 'adding' anything, it simply 'is'.
ITA did build the full thing, but only Cape Air used it. The "full thing" being shopping, schedules, inventory, booking, cancel, change, check-in, etc.
The shopping engine itself isn't trivial. ITA has the best shopping engine available.
Really similar to having a big monolith, and taking pieces out chunk by chunk rewritten to run as microservices until you're left with a new codebase with no sign of the original monolith or any of it's code.
Batch is the other aspect that takes mainframes to unbeatable status. When you run a batch job, you basically say "ok all daily OLTP is halted, we are going into a completely new architectural mode". Obtaining an exclusive lock on a table or an entire database for a single series of processes to operate on can yield insane amounts of throughput when you are finalizing the data from the day's operations.
A few years ago as a more junior developer, I will admit to being strongly "anti-mainframe" based on the principles I held at that time. Why can't we just put it all in the cloud, throw some MongoDB out there and pray it all works? After witnessing actual business cases for mainframe/batch unfold, I quickly started to change my tune. Mainframe is not for every business, but it does seem to be a tool you can reach for when someone says "everything just has to work always or someone dies", and "we have infinite money".
Due to io offloading to co-processors and a whole range of supporting cpu types you can set up a system which can handle thousands of transaction per second. Most of the application that are used on mainframe's are databases or message que's
Recently there is an increased interest in running Linux on mainframe's something which can be done since early 2000 you get the benefits of high available and secure hardware and the relative ease of Management of Linux. Another benefit is that you don't really need to train personnel in more exotic operating systems like z/OS.
Think of the number of transactions entities such banks and airlines perform a day. This article talks about this in the section "what is a mainframe today":
https://www.nanalyze.com/2017/10/ibm-mainframe-computers-tod...
I have seen a handful of mainframes on use, but I have never seen one used for something that it is good for.