At a previous job where we needed to always be up, our disaster recovery plan assumed that the us-east-1 site had been hit by a meteor (not literally, but that's how we explained it to each other to put ourselves in the mindset.)
https://aws.amazon.com/about-aws/global-infrastructure/regio...
The cli is gorgeous, the web gui is terrible.
https://en.wikipedia.org/wiki/Failure_domain
Clouds aren't magic. They require a certain amount of operational confidence in order to understand that, yes, an entire region can fall out from under you at any time and it's your responsibility to detect and deploy into an unaffected region if possible.
edit: Generally, one entire region will not fail. However, core services like STS rely on us-east-1 so it's particularly susceptible to disruption.
There are also way too many successful, public cloud-native businesses running without any semblance of a Business Continuity or DR Plan among any of their teams. I wish I could name and shame some of the more egregious cases I have seen.
Empirically? No.
But seriously, instead of making every dev team consuming an AWS product hire the extra engineers required to build a system that spans multiple failure domains — which lets be honest, companies won't — why hasn't Amazon just hired the engineers required to do it for me?
> Generally, one entire region will not fail.
Laughs in global failures.
For many use cases, it’s acceptable to shrug and blame AWS for a failure. It’s harder when your high availability solution fails independently, which they almost always do more than US-East-1
And being the only one up doesn't win as many market cred points and you'd think.
For many services, it makes more sense to make it reliable than not. For other services, it makes more sense to think about the engineering of the solution in the field.
Example: McDonald’s product images on kiosks are apparently in S3 and not cached locally. Seems like a dumb idea to me, but I wouldn’t try to build a more reliable cloud storage backend to control that risk.
FYI this example hasn’t been true for a while. STS regional endpoints are generally what you should be using these days. The “global” us-east-1 endpoint still works, and may be the default for some clients, but isn’t a requirement.
So yes, in theory much of AWS's services are probably very reliable and distributed across AZs and regions, but in practice there's likely a whole bunch of debt where one thing gets fucked up and it cascades.