Amazon EC2 outage: summary and lessons learned
blog.rightscale.com
blog.rightscale.com
I'd argue that the overall cost is much less than having all of these services in house, and in house services go down too.
But I think for the moment, they want someone to yell at, and Amazon gives the most unhelpful lack of communication, with no even remote eta, and that's unacceptable.
Given proper motivation (say 20c bonus on every dollar saved) I think we'd see this argument vanish pretty quickly.
If you listen to the wrong people (sales-guys from "Enterprise Grade" vendors), or pinch too many pennies it can easily be a disaster. It's dangerous water to tread on your own for sure.
If it's critical not to have problems if/when Amazon goes down, you have to plan for that ahead of time, just the same as anything else. It's not like throwing the "cloud" label on something makes it invincible.
As computers become increasingly integral to the daily operations of a business, the business guys are not going to have any choice but to learn some basics. We're really already at this point, but maybe after enough failures similar to those witnessed this weekend it will finally sink in.
I am amazed how many people want to blindly follow the buzzword bandwagon without obtaining even a vague notion of the technical implications first. The fact that Amazon would be controlling Amazon EC2 and that if a failure occurs at Amazon, their EC2 product may be affected, is the most blatant thing about using "Amazon EC2" or any other external service provider.
From what I gather, this is not how it was handled (I may be wrong), and for that I could not put my trust in them.
The details you want don't come until the crisis is over and the users are back online. Occasionally you may be able to get a meaningful ETA, but it really depends on the nature of the failure(s) that caused the outage. I'm glad Amazon didn't cave to the pressure and just throw a random guess out.
However, imo an executive summary that starts with "The Amazon cloud proved itself in that sufficient resources were available world-wide such that many well-prepared users could continue operating with relatively little downtime. But because Amazon’s reliability has been incredible, many users were not well-prepared leading to widespread outages. Additionally, some users got caught by unforseen failure modes rendering their failure plans ineffective." seems a little too supportive of Amazon.
I posted this yesterday, with the conjecture that it may have been a sudden sync problem. It's a good read.
The solution (I didn't reread so this is from memory) is to add random jitter to each participant's timer.
However, is there evidence to suggest that's what happened to Amazon? I can see this being a big issue in '93 with high-latency low-bandwidth links a commonplace. But we think that Amazon wasn't engineered well enough to deal with multiple orders of magnitude spikes in C&C traffic?
Thank you, though, for posting a (much needed) technical comment to this discussion.
And yes, the paper talked about randomization. It also pointed out the magnitude of randomization required was larger than expected.
At the time of writing Amazon has not yet posted a root cause analysis. I will update this section when they do. Until then, I have to make some educated guesses.
That pretty much sums it up. Well, that plus some contradictory lesson learned, such as The biggest problem was that more than one availability zone was affected, followed by, must have live replication across multiple availability zones.
Amazon's communication during this was on point. There's a line between tell me what's wrong and fix what's wrong and all of the author's suggestions on how to fix the "communication problem" are on the wrong side of that line.