A Brief History of High Availability
cockroachlabs.com
cockroachlabs.com
https://en.wikipedia.org/wiki/Tandem_Computers
"Tandem's NonStop systems use a number of independent identical processors and redundant storage devices and controllers to provide automatic high-speed "failover" in the case of a hardware or software failure. To contain the scope of failures and of corrupted data, these multi-computer systems have no shared central components, not even main memory. Conventional multi-computer systems all use shared memories and work directly on shared data objects. Instead, NonStop processors cooperate by exchanging messages across a reliable fabric, and software takes periodic snapshots for possible rollback of program memory state."
The general ideas of failover for high-availbility go back much further.
https://en.wikipedia.org/wiki/Failover
"The term "failover", although probably in use by engineers much earlier, can be found in a 1962 declassified NASA report. The term "switchover" can be found in the 1950s when describing '"Hot" and "Cold" Standby Systems', with the current meaning of immediate switchover to a running system (hot) and delayed switchover to a system that needs starting (cold). A conference proceedings from 1957 describes computer systems with both Emergency Switchover (i.e. failover) and Scheduled Failover (for maintenance)."
Also worth a mention that the well known Paxos distributed consensus algorithm was described in a 1989 paper by Leslie Lamport, and the almost identical (but less well known) predecessor Viewstamped Replication in a 1988 thesis by Brian Oki.
One of the first things that Brian Oki mentioned was Tandem NonStop, and I understand that Brian's VSR was then the first consensus protocol to deal with network partitions, where the network is not reliable, for high availability.
What was also special was that James Cowling joined me afterwards—we covered Viewstamped Replication Revisited, the 2012 revision, with a focus in both interviews on Barbara Liskov's influence on both James and Brian.
At Basho we worked hard to acknowledge our debt to distributed systems pioneers. Not everyone is ignorant of that history.
I was mainly on the periphery at Basho, started as tech evangelist, migrated to engineering to help get 2.0 out the door. John Daily.
Engineering was very talented, and the customer support team was truly top notch. Such a shame, best job I’ve ever had.
1) the more nodes, the higher probability of failure, 2) the more load, the higher probability of failure, 3) complex logic and/or nodes that aren't tested frequently have a higher probability of failure
Basically, every person who has claimed "X is easy, I do it for my personal use and it's fine with one giant node, don't need HA" is true (for them) because it's literally less likely to fail. But if they had more load they'd see more failures, and probability doesn't mean they won't encounter the random failures that are much harder to recover when there's no backup system.
HA means you're going to have more failure, which means more work to keep it running. But it's necessary work if total system failure is unacceptable. Two engines in an airplane is twice the maintenance, but it's better than the only engine conking out mid-flight.
Anybody bellow that spot is just begging for being taken out by random at the worst possible times.
I may be misreading your comment, but how is this true?
* the more complex the configuration required to provide high availability, the more likely something will go wrong and reduce availability
* the more nodes you have, the more likely one of them will fail. Whether that reduces availability typically depends on the failure mode and control layer.
For an example of the latter, with the distributed database Riak (and I imagine others as well) if a server is starting to fail (corrupt disk, bad network card, etc) you’re in a much better place if it stops working entirely. A gradual node failure is much much worse.
Really? In what sense? Hardware fails or what? The complexity usually brings everything down.
Anyway, looking forward to cockroachdb, it seems it solves some high-availability / multi-site problems.
The GP has a very digital model of computers. It's not really useful for predicting failure.
You could argue to just call the TXID "complexity" but that'd come with a wide enough definition of what complexity is that it just becomes a replacement for saying "buggy code" which isn't very useful. Most leave complexity as being more like "grossly overengineered" to leave room to avoid that.