But that was only introduced in 2010, so either the timeframe I was guessing at is wrong, or I'm misremembering what component it was. Could have been disk drives maybe.
Yes, I've seen this too, and on lowish-end servers.
Doesn't help when your mainframe's infiniband controllers shit the bed and take the whole box down (a thing which has happened to me).
In any case, it helps explain why a purse-string-holding exec would be convinced to shell out money to IBM.
I wasn't really meaning to say that a Spark cluster is as robust as a mainframe (though maybe someone's figured out some tricks), more that I could totally see some sales folks conducting the same demo using something like Spark.
You're still boned if the driver dies. I am pretty sure that the driver keeps some important state in RAM so if the node hosting it goes down you have to restart from the beginning, even if the cluster manager restarts the driver.
Meanwhile, these days you have folks running "start of the art" systems which may just roll over and die for no discernible reason. These may not easily come back up, either, if at all. That's why they have to have so many of them!
I was in college at the time and looked at job postings for mainframe programmers because I wanted to work on that. Never found one that required less than 5-10 years experience, not then and not when I've checked every couple years after that. Too bad; I quite liked the idea of working in an ecosystem that starts from "let's make this work every single time" instead of "let's make this work well enough to keep customer complaints to a dull roar".