It was physical PC clustering. So I group of networked computers appeared like 1 physical computer. While a MESO or DCOS can do this today, they don't offer durability.
On OpenVMS clusters if the machine running a job physically explodes you don't lose data. RAM and process state is network synchronized.
This allows for rolling upgrades. Your cluster can update and reboot 1 box at a time, but your apps never actually stop, TCP connections are never dropped ETC. For large corporations who've gotten used to these features they're awesome to hold onto.
One thing I wonder now: how can this work?
Is there a physical load balancer in front or something?
Is this as slow as it sounds?
At the end of the day, the windows based system was just as reliable as far as nines went and tens of times faster.
This was only just recently replaced with asp.net I understand. They got nearly 20 years out of each rewrite.
If you found a 486/100 and compared raw compute speed, the vax would likely have won for most cases, because it has 4x the registers.
I'm not actually sure about the clustering IPC performance, but you are essentially comparing hardware about 3-4 generations apart and blaming the software for the issues.. if you were comparing with 500MHz Alphas, etc, then you might have a point..
Locks on most OSes have an "outer scope" of a process or the machine.
Granted, that is as much (if not more!) a feature of the hardware architecture than the OS - but amazing still.
I've never heard of an NT box capable of that - maybe the Alpha port could do it? Anyone know?
On a more relevant note, I know x86-64 Xeon's can mark bad RAM, disk drives, and network cards, and hot-swap them all out out (and has been able to do so for > 10 years) on most motherboards (it's supported at the OS level for Windows server), not sure if it ever had full CPU failure support though.