If you have hundreds or thousands of machines, that's an indicator that you /may/ have the complexity that requires the disciplines that can come from dedicated SRE. The tough thing is conflating filling operations problems with a role named SRE, versus actually using the best practices that will help you scale and improve reliability.
If your hot standby is a$100/mo VM, it's not noticeable. If it's $5000/mo, less so.
To say nothing of scaling up and down with the load — which, of course, you only need if you are a pretty large operation.
It'll run for ten years for next to nothing.
And you also need to have diverse routing for power coming in and generator / battery room set up.
I think the main issue is that the cloud providers don't publish much about outages that don't affect the end-user. I mean a failed hard drive happens all the time, but S3 is never affected by that.