Summary of the AWS Service Event in the Sydney Region
aws.amazon.com
aws.amazon.com
Whereas Google was recently down for less than 18 minutes. A VP at Google sent an email advising all affected customers, posted continuous updates to their status page, sent a further apology email at the conclusion, posted a service credit exceeding the SLA to all customers in the zone (without forcing customers to chase this themselves with billing) and lastly wrote one of the most well written post mortems I've ever seen. AWS has much to learn from Google about how to handle outages properly.
Compared to the sole region ap-southeast-2 (for you) in AWS's case. (though it was a longer outage).
That being said, I 1000% agree that Google cloud platform's response to fix issues, postmortems and actions to make things right are top notch.
Can you explain this a little more? Amazon says this only affected one AZ, and they specifically note:
For this event, customers that were running their applications across multiple
Availability Zones in the Region were able to maintain availability throughout
the event.Apart from one internal project which mistakenly had all it's app server instances in -2b (ooops!) - all my production mobile app backends are spread across the 3 Sydney AZs. That's a few dozen EC2 app servers across about 15 projects.
My monitoring reported a worst case of 57 seconds of degraded connectivity - which was an instance in -2b going offline and the ELB not taking it out of the rotation very quickly, the app running on that had interruption, but only while waiting for the timeouts. Crashlytics and GA crash reporting didn't bat an eyelid... I had under 70 users active at the time, 1/3rd of them may have seen a minute or less of loading spinner if they'd fired of a UI blocking api call during those 57 seconds. I'm not looking _super_ closely, but nothing I'm monitoring apart from EC2 - like RDS, S3, ELB, SNS - showed _any_ glitches (I'd _probably_ have caught even single digit second problems for _some_ of that...)
I'm actually quite happy with how everything went - we don't go to any particular heroic lengths to ensure HA or uptime, we just follow recommended best practice, and at least in this outage, that worked out fine for us (except for that project where all the app servers were in -2b, and I'm happy to wear that as our fuckup)
While the numbers are nice, if some outages only impact a subset of customers and your monitoring accounts aren't one of them it's hard to determine how good your monitoring data really is. If he was impacted by a 12 hour outage and you only show ~2 hours that's a really significant difference.
I guess it really depends on how you monitor and how comprehensive it is. Do you monitor from multiple ISPs on different network paths in multiple regions/countries? Do you monitor each of the services under different load conditions and monitor multiple accounts? Sometimes "up" only tells part of the story.
Your monitoring sounds really comprehensive. That's a very cool way to advertise the service it's built on. Do you monitor service providers for outages that reduce capacity or increase latencies but otherwise the service is "up"?
(Mutter, mutter, … something about Americans and their timezones … and northern hemispherians and their seasons …)
Sun July 5th
3:25 PM -- Initial Power outage
4:42 PM -- Instance launching in unaffected AZ's restored
4:46 PM -- Power Restored
6:00 PM -- 80% of instances recovered
7:49 PM -- DNS recovered
Mon July 6th
1:00 AM -- Almost all instances recovered* http://www.smh.com.au/national/australias-wild-weather-sydne...
* http://www.abc.net.au/news/2016-06-07/sydney-weather-storm-d...
* http://www.sbs.com.au/news/gallery/pictures-wild-weather-sav...
It is false to assume that the state of the electrical supply is either on or off. This may come as a surprise, but not to me. In 2008, Eskom (South Africa's electricity suppliers) experienced similar faults. The mains supply voltage is 220v here. At one point, some devices started to fail in my house, and others, such as lights, continued to work, but significantly dimmer. We measured 180v at the plugs. There were similar outages in my area last year, where an outright cut-off was preceded by voltage drops. This outage is interesting because it is an example of a bug owing to false assumptions!
There have also been incidences where certain cables have been stolen [1] and that has caused the opposite: voltage spikes.
[1] I couldn't tell you which, or what kind, but I remember it has something to do with "the neutral"
E.g. in this case, in normal operation, power from the utility power grid spins a flywheel. When the grid fails, the flywheel provides a holdover until Amazon's diesel generators can start.
But in this failure the voltage from the grid sagged, rather than going away completely. The breaker isolating the flywheel from the grid didn't open quickly enough. So power from the flywheel was sent out to the grid. It didn't succeed in powering the grid for very long. Oops.
Agreed. I'm still put off by the fact that ELBs specifically can not handle a sudden spike in traffic orders of magnitude higher than the previous rate. They fall on their face, bad. If you expect a spike like that to happen, you literally have to submit a ticket and ask them to pre-warm your ELB...
I am reminded of 'The Event' from That Mitchell and Webb Look [0]. We don't talk about The Event.