I suspect this is more about not having to pay IBM for the hardware and contracts anymore
I suspect this is more about not having to pay IBM for the hardware and contracts anymore
I expected a drop in performance when a particular contract I worked with asked us to see if we couldn't get everything running on a VM instead of the IBM iron. Like you suspect, that was about not paying for the hardware anymore because just the price of electricity was a noticeable impact on the bank's overall budget. It was an R&D project to see if we could avoid that.
What we put together was a QEMU instance, running atop 3 very very cheap commodity servers in duplication, so that if one went down it would switch over without downtime. We chose the cheap servers, because we already had them hanging around. (Probably around $1000 in hardware all up, today.)
We did not see a drop in performance or reliance. But we did see an increase in performance. As in, ~30%. I avoided suggesting this is always the case, but it is a significant possibility when changing. As far as I know, the bank in question is now running a similar setup everywhere they used to have a mainframe.
I am very curious what you mean by "in duplication". Do you mean "3 copies of production and a router" (probably not) - or are you saying you did something like lockstep for QEMU?
Incidentally I just found https://wiki.qemu.org/Features/COLO, which looks like it may be being successfully used privately in one or two places (looking at the email addresses).
I'd probably use COLO if I was to do it again today.
[0] https://kashyapc.fedorapeople.org/QEMU-Docs/_build/html/docs...
I wonder what sorts of scenarios would benefit from using this instead of any of the dozen distributed database architectures it seems are out there.
Incidentally while poking through https://www.qemu.org/docs/master/system/invocation.html I noticed that there are some COLO-related options in there, which is a bit exciting.
It's there for safety reasons for automotive. (see https://blogs.nvidia.com/blog/2020/05/20/xavier-achieves-ind... )
The first widely available ARM cores providing it are fairly new (at least in the automotive domain).
[0] https://developer.arm.com/ip-products/processors/cortex-a/co...
As a person who worked with Xavier for quite a while, dual-core lockstep is supported. Nvidia uses their own CPU cores, not Arm's.
See: https://docs.nvidia.com/jetson/archives/l4t-archived/l4t-323... for how to enable it.
> enable_ccplex_lock_step: Boolean; enables or disables CCPLEX dual-core lock step.
Heavily vectorised code with modern instruction sets can also be more performant than a lot of the older compute chips, but that requires more rewriting of code, and someone who intricately understands both the code and the math. Which makes modernising much more expensive.
The VM approach is a simpler way to get you most of the way there, but replacing heavy compute stacks is usually going to require a decent bit of investment.