S3 and High Availability
blog.movableink.com
blog.movableink.com
But the guidance we've been given by Amazon is that this is the purpose of availability zones, not regions. Regions are more appropriate for fighting the speed of light (i.e. locating your site closer to your users). As an illustration of this, Amazon told us that the US version of amazon.com runs in a single region.
Incidentally, the other interesting take away from that meeting was to avoid using autoscaling to respond to failures. This is because provisioning instances can fail when there's heavy demand and that's frequently the case when Amazon is experiencing outages in other regions and availability zones. Instead, we've been urged to provision 150% of what we need (50% in each AZ) so that if any one AZ goes down, we can still handle all our traffic. Where autoscaling works well is in responding to spikes in our own need rather than situations where many Amazon customers will have need.
Sorry for the digression, but I found that consultation interesting and it's clear that others have the same misconceptions that I had before learning the thought process behind the building blocks that AWS gives us.
The S3 issue affected the entire region. A month back a BGP issue affected every one of our AZs in us-east-1. It wasn't even Amazon's fault, but a multi-AZ strategy would have done nothing for us.
It's kind of like Maslow's hierarchy of availability needs: at the bottom you deal with hardware failures, then networking failures, then regional configuration failures, and finally homogeneity issues where all of your machines fail at once because of a bug in hardware raid controllers or a leap second smear issue. It all depends on how much resilience you need and are willing to pay for.
Devil's advocate, though: if you're concerned about what to do in the event of an AWS regional failure, given how rare an event that is, then you've probably outgrown AWS.
(For most small-to-medium sized startups, I'd advocate setting up statuspage.io and keeping your users informed if you're single homed to an AWS region and that region experiences a catastrophic failure. The math on "how much money you'll lose from the outage" vs. "how much you need to spend implementing proper DR, better than what AWS has in place to keep a region up and running" isn't even close, assuming, say, 1 8 hour regional outage every ~2 years.)
Where I work now (a small web-based software-as-a-service) an 8 hour outage would be catastrophic, and could easily kill the business. The switching cost for our niche is small, so one bad day could cost us 20% of our clients. We're not running with a 20% profit yet, so at best it'd mean an immediate layoff or across-the-board temporary paycut.
Luckily, because we're small we can run on a single LAMP server. We're working on making it so that we can migrate that to any region in EC2 with a single command, as well as making sure we can switch to a different dedicated hosting provider with minimal time.
Sound goods? (Woh, as long as Cloudflare doesn't have any SSL issue :D)
Historically this applied to all regions except US Standard, but now that too supports it if you go through the VA instead of global endpoint.