* Azure: 12 user-facing outages this year so far. source: https://azure.microsoft.com/en-us/status/history/
* Amazon: They don't make their status history public it seems, but I see 5 easily searchable outages from last year. I would venture to bet their actual outage history is not any better than Azure or Google.
* Google: 23 from this year so far source: https://status.cloud.google.com/summary
These numbers are from when I counted in March 2019, so ~ 3 months, their outages are in double digits. That's TERRIBLE.
Obviously the above numbers are not across _all_ regions, they are just a total count across all regions.
If I really need perfect availability (I probably don't), then I can leverage the multiple regions and build a failovers strategy.
I do not expect global outages which affect multiple regions at the same time. That's just amateur hour when it is a common occurrence.
There are many ways to handle outages.
I'm just saying if you are averaging 4+ outages PER MONTH across all of your services, you are probably doing something wrong. These providers are on schedule to pass 50 outages this year!
All the places I've ever worked, if we can count our customer facing outages per year past the number of planets humans live on(1) we consider ourselves a total failure that year and re-work our planning. We haven't had to re-work our plans in ~ 5 years. Our longest outage in the past 5 years? 1 hr long. Granted I'm not the scale of $CLOUD, but there is zero reason they can't get serious about uptime, they just don't bother. People are still flocking to $CLOUD in droves, despite their crappy uptimes.
It's not rocket science to make stuff work, you just have to get Boring(tm).
Where Boring here means using standard, boring, well understood technology that everyone knows and understands and not $NEWHOTNESS. $NEWHOTNESS is guaranteed to break in new and exciting ways.
I'll be interested in the post-mortem from Azure on this one.
(I had to, see username!)
edit: seriously,-3 ? it was a joke.
Well, that's interesting. We occasionally see getaddrinfo() calls fail claiming domains that we know exist at the failure time (b/c the records are completely static) don't exist. (We've not got a reproducible case for this yet, and it's incredibly rare for any given VM/service. But across our fleet, it crops up fairly regularly.)
That could be whatever resolvers you're hitting failing rather than an issue with Route 53 authoritative nameservers, though. The resolving DNS servers in EC2 are not actually part of Route 53, for example.
But, the stuff that hits this problem the most often is of the quality level that I wouldn't find that terribly surprising. Seems AWS "documents" this as,
> The number of DNS queries per second supported by the Amazon-provided DNS server varies by the type of query, the size of response, and the protocol in use.
How specific.
That said the most common cause of authoritative nxdomain is if youre adding/deleting records and querying them before propagation is complete. You may want to log/poll your rrset change status separately to correlate.
The other is that depending on networks intermediate dns tampering happens all the time. Qname, rname, rtype, all get modified. Responses and queries are duplicated, intercepted, and manipulated. Some good research out of dns oarc and a dude out of australia (iirc).
There was the status page, but S3 being down in us-east-1 didn't effect S3 in ap-southeast-1, etc. The big DynamoDB outage a few years back also was limited to us-east-1