Most companies would have said "the outage was caused by a system being accidentally misconfigured" or even less; Amazon, in contrast, admitted (a) that a specific person was identified who made the mistake, (b) that he had access to the systems in question because a process was being run manually which should have been automated, and (c) that if there hadn't been an error in the access control rules, he wouldn't have been able to make that mistake.
I often wish that Amazon was more open about their internal systems; post-mortems are the one time when I'm never wondering what they're not telling me.
I course, I am small potatoes compared to Netflix being down (while Amazon instant video wasn't.)
Amazon is populated by jackasses apparently and Heroku isn't much better. Why the f*(k doesn't Heroku have some kind of fall-over or something to protect clients from the inevitable AWS-East failures?
Buh, bye Heroku and AWS. Hello Bluebox. I run other apps on Bluebox and have never had any downtime except for some mild nonsense due to my own damned ignorance. Luckily the Blueboxers saved me from myself. AWS needs to get their shit together.
That said, I understand it's pretty damn frustrating having downtime during peak periods :)
Make smarter decisions. If being up on December 24 was really that important to you, you'd have had backup hosting in place with the ability to quickly fail over to it. You'll start being a better engineer when you quit blaming some poor bastard's bad luck for your failures and learn the real lesson from your downtime: you fucked up by not having high-availability hosting to meet your claimed high-availability needs.
The worst thing Amazon could do is fire this guy. They've just paid a lot of money to have him learn from his mistake. They'd be fools to throw that hard-earned experience away.
People can't own up to the fact that customers don't really care if it's Amazon AWS or Heroku or some other platform that failed -- if your service is down, its your fault!
I don't, necessarily, have a problem with saying, without animosity, "we're down because our hosting provider is down". For many small businesses, being down when your hosting provider is down is an acceptable tradeoff -- your business simply isn't critical enough for it to be worth the costs of maintaining fail-over redundancy in case of primary hosting loss.
That's totally fine. Most people probably don't want to pay the higher cost of having every service they use having the ability to survive their main host going down. A couple 9s of uptime is more than enough for the vast majority of businesses.
What I have a problem is with people who were behaving as though they were fine with that tradeoff screaming about how they needed to be up when it comes time to pay in a little downtime for those cheaper hosting costs. Grow up and accept the tradeoffs you choose to make.
I'm not sure who is to blame at this point -- is it the fault of the individual for not understanding uptime? is it the fault of the service for not articulating clearly what it means? Is it the fault of the industry for inculcating an unjustified sense of entitlement?
I'm inclined to lean towards the third option. It ought to be obvious that there is no such thing as a perfect host, that in the limit your probability of being down at some point goes to 100% as your time hosted with a company increases.
There's just a lot of entitlement and sloppy thinking out there. People believe that their $30 a month hosting ought to be magically invulnerable in ways it by the simple laws of physics cannot be. If you didn't explicitly choose to have multiple geographically separate hosts with a fail-over plan between them, then you were at the very least implicitly accepting that you were going to be down some of the time when (not if) your only host goes down. That's a fine and ok thing to choose, but I don't think you get to be mad at your host because you didn't clue in to the choice you were making.