I think one decent case would be to keep yourself honest.
Linux started out a 80386 only, but someone ported it to DEC Alpha. By deciding to run on both 32- and 64-bit platforms in the 1990s, it kept the kernel developers agnostic. Then when the x86 world went 64-bit with amd64, there was probably a lot less 'cleanly up' to do for it to be ported compared to if it had been focused on pure-x86 (see also SPARC and endianness). It's said that NetBSD has a very clean code base because of it's reputed high portability.
Similarly Solaris was very scalable. Partly because in 1996 Sun released the SPARCcenter 2000, which could handle 20 CPU sockets. That was a lot, and I'm guessing not many folks bought one, but they had to deal with it in an official capacity, as time when on more CPUs (and cores) became prevalent Solaris was good to go as that situation became more mainstream.
If you develop for the odd ball situations and corner cases, it may force you to be less lazy as a programmer.
My main gripe with the various 'E-series' systems (IIRC, we had some E3x00s) was boot-up time, especially the RAM check on power up.
I was in an academic environment, and so every so often we had to power down our lab for electrical/fire code inspection power outages by Facilities. The first time we rebooted one (with 4GB of RAM?) it kept booting and booting and booting and booting and booting and booting. 17 minutes later we got the login prompt on the console.
We put the 17 minutes as a note in our run book so that we knew not to worry if took 'forever' and to move onto the next step. Otherwise we'd start freaking out about something being "broken" with the system(s).
(This was in the 2001-2001 timeframe.)
I don't think we ever got the OC-12 ATM cards really stable...fortunately, networking went a different direction and they had a mercifully short lifespan.
Replacing a mainframe with a VM and not losing performance is a huge benefit.
But the software that used to run on that mainframe is now easier to modify, and there's a 10 year backlog of critical bugs that need to be fixed but we were always too scared of breaking something because we couldn't test it - now we have a VM. We can spin up a second one.
I suspect this is more about not having to pay IBM for the hardware and contracts anymore
I expected a drop in performance when a particular contract I worked with asked us to see if we couldn't get everything running on a VM instead of the IBM iron. Like you suspect, that was about not paying for the hardware anymore because just the price of electricity was a noticeable impact on the bank's overall budget. It was an R&D project to see if we could avoid that.
What we put together was a QEMU instance, running atop 3 very very cheap commodity servers in duplication, so that if one went down it would switch over without downtime. We chose the cheap servers, because we already had them hanging around. (Probably around $1000 in hardware all up, today.)
We did not see a drop in performance or reliance. But we did see an increase in performance. As in, ~30%. I avoided suggesting this is always the case, but it is a significant possibility when changing. As far as I know, the bank in question is now running a similar setup everywhere they used to have a mainframe.
Heavily vectorised code with modern instruction sets can also be more performant than a lot of the older compute chips, but that requires more rewriting of code, and someone who intricately understands both the code and the math. Which makes modernising much more expensive.
The VM approach is a simpler way to get you most of the way there, but replacing heavy compute stacks is usually going to require a decent bit of investment.
I am very curious what you mean by "in duplication". Do you mean "3 copies of production and a router" (probably not) - or are you saying you did something like lockstep for QEMU?
Incidentally I just found https://wiki.qemu.org/Features/COLO, which looks like it may be being successfully used privately in one or two places (looking at the email addresses).
I'd probably use COLO if I was to do it again today.
[0] https://kashyapc.fedorapeople.org/QEMU-Docs/_build/html/docs...
I wonder what sorts of scenarios would benefit from using this instead of any of the dozen distributed database architectures it seems are out there.
Incidentally while poking through https://www.qemu.org/docs/master/system/invocation.html I noticed that there are some COLO-related options in there, which is a bit exciting.
It's there for safety reasons for automotive. (see https://blogs.nvidia.com/blog/2020/05/20/xavier-achieves-ind... )
The first widely available ARM cores providing it are fairly new (at least in the automotive domain).
[0] https://developer.arm.com/ip-products/processors/cortex-a/co...
As a person who worked with Xavier for quite a while, dual-core lockstep is supported. Nvidia uses their own CPU cores, not Arm's.
See: https://docs.nvidia.com/jetson/archives/l4t-archived/l4t-323... for how to enable it.
> enable_ccplex_lock_step: Boolean; enables or disables CCPLEX dual-core lock step.