Tandem Guardian 90: A Distributed Operating System [pdf] (1990)
hpl.hp.com
hpl.hp.com
A problem with many non-stop setups, including fault tolerant distributed computing, is that sales people loved to pull cards on the mini, or switch fabric, to show how the machine kept going. I heard stories of some back-end engineers who were continually receiving CARDFAIL events, and sending field parts out to the display room and worried about a sustaining build/reliability problem until they worked out what was going on.