So in my now close to 30 years of experience later, working on systems expected to run 24/7/365, I can't tell you what percentage of failures and outages were caused by the redundancy/fail-over/etc software and hardware layered into a system, but its definitely a very significant part of the problem. Frequently because its just that a "layer" someone bolted on, even if that layer was some software engineer designing for "web scale". All the edge cases are frequently not completely thought out, and when you hit one its a lot harder to recover, than simply restarting a single service running on a single server.
I’ve had massive problems with fancy (HPE, IIRC, but it wasn’t called HPE then) raid arrays caused by the controller being a POS. Using plain mdraid would have been far more reliable.
Sometimes I wonder if all the effort people put into protocols that run Paxos and Raft could be better spent running a small number of coordinator servers with manual failover. I’m suspicious that, in many use cases, leader elections cause failures more often than the leaders themselves fail.
There has to be a smart reply to that, like: 100% of pilots of single engine planes don’t make it home after a single engine failure. 100% of pilots of dual engine planes do make it home after a single engine failure.
pushes glasses
Sounds like someone didn't understand redundancy...
So the odds are you don't. Keep it very simple, very small. I can't highlight this enough. A few boxes will get you very far. If you ever actually hit the limits, throw a big party to celebrate success and only then start to think about making your infrastructure a bit (just a bit) more complex.