http://www.slideshare.net/twilio/highavailability-infrastruc...
http://www.twilio.com/engineering/2011/04/22/why-twilio-wasn...
It's strategy as opposed to how-to but the principles apply.
http://www.slideshare.net/twilio/highavailability-infrastruc...
http://www.twilio.com/engineering/2011/04/22/why-twilio-wasn...
It's strategy as opposed to how-to but the principles apply.
http://techblog.netflix.com/2011/07/netflix-simian-army.html
They've even released the "Chaos Monkey" open source: http://techblog.netflix.com/2012/07/chaos-monkey-released-in...
There's pretty much no way to architect around that one as an AWS user (apart from going fully multi-cloud, but "nobody" actually does that, at least at scale), and I'm kind of shocked that those bits of AWS are still not robust against "single AZ outages", given that they're involved in pretty much every one of these incidents and make them affect people on the entire cloud...
Pirate Bay might disagree with that sentence: http://torrentfreak.com/pirate-bay-moves-to-the-cloud-become...
But regardless it's not like all of EC2 went down just one or two AZs. So why couldn't traffic be migrated transprently to other AZs/regions ?