Designing AWS Architecture to Withstand Outages
madeiracloud.com
madeiracloud.com
If your application is architected to use SimpleDB, DynamoDB, SQS, or even RDS, these simple "You should have been using Route 53 and multiple regions" articles get increasingly frustrating. Most applications simply aren't elementary enough to fit into that boilerplate structure, and getting around that fact either requires switching away from Amazon's managed database services (lots of money) or writing synchronizing scripts that play nicely with your stack and launching them on additional servers (lots of money and lots of resources).
While ideally I'd like to see Amazon release some sort of feature for multi-region sync, it would be interesting to see how the tried-and-true multi-region businesses have approached this problem.
I think Google Compute Engine does exactly that, with a direct connection between regions so the entire system works as one network. I'm not sure if that will be the case outside the US, though.
For the SQL datbases they provide a sync service that can replicate your database automatically between different regions (or even a local database).
I have no clue how reliable all of this is, but at least in theory it looks as if it might be quite a bit easier to do full geo replication across different regions on Azure.
Anyone with actual experience?
The diagrams show two separate masters in different regions with their own slaves, yet DNS is essentially randomly delivering users to each "half". There is no mention of how this is handled, or even that you would have to consider it.
Yes I know, business continuity and stuff. Still, just doesn't feel right somehow.
If you run a service where the worst case scenario for your site being down for an hour on a Thursday afternoon is "millions of dollars are lost" or "our high profile customers go out of business in a way that's directly traceable back to us" then yes, you need all 24 of those things.
If, on the other hand, the worst case scenario for your site being down for an hour on a Thursday afternoon is "some of our customers have to manually post to Facebook so that their friends know that they've been for a run", you can probably shave about 21 nodes off that diagram.
You can then pay for three really cheap VPS's and load balancers around the world using something like Linode (or a much less reliable VPS, doesn't really matter) and replicate your data occasionally to the VPS. When your AWS instance goes down (because you foolishly rely on the same Virginia datacenter that has more breakdowns than Lindsay Lohan) you cut over to the hot spare VPS and rate-limit your incoming requests until AWS comes back.
You end up paying for 6(ish) services, still relying mostly upon AWS but with a tiny DR site you can use during emergencies.
Don't bother with AZs, they've proven not to be an independent unit of availability.
Does anyone have any experience with it?
We're still in early days but we did a Show HN a little while ago: http://news.ycombinator.com/item?id=3808031