So in my now close to 30 years of experience later, working on systems expected to run 24/7/365, I can't tell you what percentage of failures and outages were caused by the redundancy/fail-over/etc software and hardware layered into a system, but its definitely a very significant part of the problem. Frequently because its just that a "layer" someone bolted on, even if that layer was some software engineer designing for "web scale". All the edge cases are frequently not completely thought out, and when you hit one its a lot harder to recover, than simply restarting a single service running on a single server.