Could be a different incident and a different machine, though. I'm sure this story happened more than once.
The infamous machine did go through repairs and part swaps many times, as you could see from its long and troubled hwops history.
The worst machines were the zombies with NICs bad enough to break Stubby RPCs, but still passing heartbeat checks. Or breaking connections only when (re)using specific ports. Fun times!
Regarding this system: the motherboard was never swapped?
> engineers habitually run batch jobs with more replicas than there are machines
Idly curious, how do I parse parse this? It sounds like the same jobs are replicated to multiple machines as a sort of asynchronous, eventually-consistent lockstep arrangement?