Doesn't the region going down mean that _all_ its AZs have gone down? Or is my mental model of this incorrect?
A region is a networking paradigm. An AZ is a group of 2-6 data centers in the same city more or less.
If a region goes down or is otherwise impacted, its AZs are unavailable or similar.
If an AZ goes down, your VMs in said centers are disrupted in the most direct sense.
It's the difference between loss of service and actual data loss.
So when failures happen in that region (and they happen more commonly than others due to age, scale, complexity) then they can be globally impacting.
Availability zones are supposed to be another fault boundary, and things are generally pretty solid, but every so often problems spill over when they shouldn't.
The general impression I get is that us-east-1's issues tend to stem from it being singularly huge.
(Source: Work at AWS.)
Literally all our AWS resources are in EU/UK regions - and they all continued functioning just fine - but we couldn't sign in to our AWS console to manage said resources.
Thankfully the outage didn't impact our production systems at all, but our inability to access said console was quite alarming to say the least.
It would probably be clearer that they exist if the console redirected to the regional URL when you switched regions.
STS, S3, etc have regional endpoints too that have continued to work when us-east-1 has been broken in the past and the various AWS clients can be configured to use them, which they also sadly don't tend to do by default.
These big outages are noteworthy because they _do_ affect people who correctly architected for reliability — and they're pretty rare. This one didn't affect one of my big sites at all; the other was affected by the S3 / Fargate issues but the last time that happened was 2017.
That certainly could be better but so far it hasn't been enough to be worth the massive cost increase of using multiple providers, especially if you can have some basic functionality provided by a CDN when the origin is down (true for the kinds of projects I work on). GCP and Azure have had their share of extended outages, too, so most of the major providers tend to be careful to cast stones about reliability, and it's _much_ better than the median IT department can offer.
“If you're having SLA problems I feel bad for you son I got two 9 problems cuz of us-east-1”
If you wanted to deploy something similar, like Cassandra across AZs, or even regions you're welcome to do that. But now you're on the hook for the availability of the system. Are you going to get higher availability running your own Cassandra implementation than the DynamoDB team? Maybe. DynamoDB had a pretty big outage in 2015 I think. But that's a lot more work than just using DynamoDB IMO.
this is both more and less true than you might think. for most regional endpoints teams leverage load balancers that are scoped zonally, such that ip0 will point at instances in zone a, ip1 will point at instances in zone b, and so on. Similarly, teams who operate "regional" endpoints will generally deploy "zonal" environments, such that in the event of a bad code deploy they can fail away that zone for customers.
that being said, these mitigations still don't stop regional poison pills or otherwise from infecting other AZs unless the service is architected to zonally internally.
[1] https://www.reuters.com/article/us-france-ovh-fire-idUSKBN2B...
[0]: https://aws.amazon.com/blogs/aws/in-the-works-aws-canada-wes...
Our SOP is to cut over to a second region the moment we see any AZ-level shenanigans. We've been burned too often.
Usually there would be high network error rates which were usually enough to make RDS Postgres fail over if it was in the impacted AZ
The only real "outage" was DNS having extremely high error rates in a single us-east-1 AZ to the point most things there were barely working
Lack of instance capacity, especially spot, especially for the NVMe types was common of CI (it used ASGs for builder nodes). It'd be pretty common for a single AZ to run out of spot instance types--especially the NVMe ([a-z]#d types)
My system is latency and downtime tolerant, but I’m thinking I should move all my Kafka processing over to us-west-2