In the flip side:
> the problem with a single system is that mundane things cause unavailability
Is multiplied by each "system" you add. The basic failures are relative easy to deal and understand, the ones introduced by more complex system who knows?
Probably the most important step that I miss:
> How long would it take to procure a new server and restore from backup?
.. and applied to any complex extra you have.
For single system is viable (with something like nix) to go off for maybe half hour for what I see around, most of the time in procuring another machine (that what people do is to Bring any other machine it can not go to amazon and buy!) and restoring the backup.
Of course, I factoring that downtime is not "seconds or minutes" here, but neither I think many can do like that
P.D: all