http://www.slideshare.net/twilio/highavailability-infrastruc...
http://www.twilio.com/engineering/2011/04/22/why-twilio-wasn...
It's strategy as opposed to how-to but the principles apply.
http://techblog.netflix.com/2011/07/netflix-simian-army.html
They've even released the "Chaos Monkey" open source: http://techblog.netflix.com/2012/07/chaos-monkey-released-in...
There's pretty much no way to architect around that one as an AWS user (apart from going fully multi-cloud, but "nobody" actually does that, at least at scale), and I'm kind of shocked that those bits of AWS are still not robust against "single AZ outages", given that they're involved in pretty much every one of these incidents and make them affect people on the entire cloud...
Pirate Bay might disagree with that sentence: http://torrentfreak.com/pirate-bay-moves-to-the-cloud-become...
But regardless it's not like all of EC2 went down just one or two AZs. So why couldn't traffic be migrated transprently to other AZs/regions ?
After this issue is over I can give a longer answer. In short, we've just evacuated the affected zone and are mostly recovered.
And +1 for the slideshare page.
Their techblog is also worth following: http://techblog.netflix.com/
Since you've mostly recovered, how did your system do? Are there side-cases that Chaos Gorilla didn't touch?
Currently, Netflix uses a service called "Chaos Monkey" to simulate service failure. Basically, Chaos Monkey is a service that kills other services. We run this service because we want engineering teams to be used to a constant level of failure in the cloud. Services should automatically recover without any manual intervention. We don't however, simulate what happens when an entire AZ goes down and therefore we haven't engineered our systems to automatically deal with those sorts of failures. Internally we are having discussions about doing that and people are already starting to call this service "Chaos Gorilla"."
I would argue that none of the common full stack frameworks that startups use are fault tolerant enough for AWS. Most of them have multiple failure points that can quickly bring down entire apps.