Amateur hour at AWS
blog.opbeat.com
blog.opbeat.com
They got it fixed within 48 hours, globally, which, if you ask me, is incredible at their scale.
I would hardly describe anything AWS does as amateur. But maybe that's just me.
Sorry their updates weren't to your liking but they were responding and posting bulletins the whole time and again: they solved the issue very quickly given the number of clients they support.
It's pretty fucking professional to update the infrastructure that runs half the internet in under 48 hours with no issues. But again, communication can be a problem when you have as many customers as they do.
OP raised some legitimate concerns, but his credibility was undercut by attacking Amazon and calling them names. Ironically, his post was a much more amateur move as his concerns would likely be taken more seriously if he had stuck to the issues and not resorted to name-calling. The essence of professionalism is sticking to the issues at hand and not being sidetracked by extraneous factors.
Folks, this is an account created 19 hours ago making half an apology and then rationalizing not listening to customers because they solved a "big" problem quickly. That's an ad hominum argument and as a result, a big tell on intent.
From my perspective, pushing out a new SSL build to a bunch of load balancers in a highly automated network like AWS is probably, by this point, a trivial task. Actually listening to the customer and responding decently is MUCH harder. Clearly it could be done better, which is the point of the post.
Rise above getting offended/scared about being called "amateurs" and start talking more about what goes on in that creepy black box that is AWS. You owe the world that much, at the very least.
To clarify: I'm a long time HN reader that finally got around to making an account (and certainly not for the express purpose of defending Amazon). However, I did want to call the author out on writing a terribly unfair knee-jerk, heat-of-the-moment indictment of AWS (this type of thing is unfortunately all-too-common in the tech community: actual amateurs writing as if they are a central authority about subjects that they have something approaching 0 understanding of. For example: the multitude of complex engineering and PR challenges a service provider like AWS faces during something like the Great OpenSSL Exploit of 2014). What I'm trying to say is: cut them some slack. Their response seemed perfectly reasonable to me.
Hope this helps.
The way to respond to an asinine comment, if you must respond, is to politely refute it.
In retrospect, what I should have done is called out the blaming statements you made in your first post. That's what brought me to action and caused me to write my response the way I did. I should know better than trying to rationalize with someone who is in dissonance. BTW, narrowrail called you out below for this blaming statement here. Pay attention - people are giving you feedback. Take it or leave it.
Vote down all my comments if that makes you feel better. Karma is meant to burn. It's also a tell that this story dropped off the main page and I'm still getting downvotes on my comment. AWS koolaid much?
Oh, and FWIW, I am a super sleuth. A super sleuth of human behavior and emotional response. I also watch what I say about others, trying not to blame and indicate opinion where needed. That's why I said your behaviors were a 'tell on intent'. I have no idea who you are or why you created an account just to comment on this story, but I guarantee there is more to it than what meets the eye.
These responses I always see on HN when there is an AWS issue always show me how disconnected many of the commenters are from reality, or from ever being involved in a huge infrastructure.
Sure the AWS status page doesn't have a hip web 2.0 AJAX backed d3.js powered cool looking status page. Yes they don't update it every 3 minutes with new info, but many (most) of the problems that one off customers see do not reach a threshold that would ever effect enough customers to make it into a dashboard post. I do think they need to speed up their status updates, but these posts need to get OK'd by a decent number of people before they get thrown up.
There are usually multiple ELB instances living on every rack of every datacenter in every AZ in every region of AWS. Relaunching / patching hundreds of thousands of instances in 48 hours with minimial disruption to customers, is a lot harder than you think.
My primary point agrees with your second paragraph, which is that they could do better on the status updates. Unfortunately this has been going on for YEARS at AWS, so it's worth ratcheting up the tone when talking about it. It's important, and they need to fix it.
It's a nightmare to know when your problem is due to your infrastructure or if their's a bigger scale issue at AWS cause they never talk about it...
Sadly this page is not updated often enough. I mean, when there's a known issue on one of the aws service this page display a little "i" icon, which is barely visible.
And when, you encounter some problem with your AWS stuff that clearly come from their side, if the problem it's not wide, they just say nothing. At that point you can search yourself for hour to be sure it's not your responsibility, and after-woods, you just wait, blind.
In this case, they probably didn't want to be too explicit about the details of patching tens of thousands of machines while the remediation was still ongoing.
I do agree that it's unexpectedly hard to find a link to the "security notice" page anywhere.
Also, how'd you find that link? If you happened to just have it laying around, that's fine but it would be better if they had these things linked somewhere customers can find them when new ones are posted (like a page covering service post-mortems) and the timeline is missing little details like the year when the outage happened and a point of contact (it's signed by "the aws team") if you have questions.
That's a great point. I found it by googling "AWS post mortem", but I only knew it existed because I had been linked to this page before.
Despite not putting time stamps on their communications (which you seem really really upset about), they fixed everything for everybody in like a single day. You, their customer (and better still, me, their customer) didn't have to lift a finger.
This is exactly why we farm out infrastructure to companies like Amazon. They have a whole squad of smart people standing around leaning against a post all day every day waiting for something like this to happen so they can jump in and fix everything for us.
My little one-player business has no such team o' dudes on standby. If it weren't for the fact that Amazon is cleaning up for me, I'd be two days into having a really bad time and getting no productive work done.
So yeah, you'd be still having a bad day. Because you have no team 'o dudes on standby, so your already busy staff are chaffing at the bit to get the certificates sorted out, and potentially not being able to restore service and constantly checking to see Amazon has completed their patching. i.e. little to no productive work is being done.
https://aws.amazon.com/security/security-bulletins/
The author suggests that Amazon only made one post, updating it throughout the process. But at the link above you will see four posts regarding this issue, with the first one having been updated once to add information. No, they are not timestamped, but they are dated.
While it may be fair to criticize AWS for its customer communications during this process, I'm fairly certain that if he had a backstage pass to such a comprehensive process of remediating thousands of production systems with zero downtime, he would perhaps find the team somewhat less... amateurish.
I'd even argue that it's not a good idea to advertise, "these servers are vulnerable to this attack".
AWS is massive and organizing that kind of update by an army of engineers isn't easy.
You received a non-personalized message because AWS support probably received tens of thousands of irate customers demanding that their systems be patched immediately. For some reason they weren't equipped to handle that kind of update but I'm sure they will learn from this and hopefully next time the response will be faster, if possible.
* "However, due to nobody outside of AWS knowing exactly how ELB works, this could just mean that the machines currently responding to the requests are patched, but in the next request, it could hit an unpatched ELB machine."
* "We then wrote to support to hear if it was now safe to re-key the certificates, but did not hear from them for hours."
Summary: You re-key your certificates, thinking you are all good. Now an attacker hits a non-patched ELB, exploits the issue and gets your new keys.
and then update your certs again just to be safe.
If you absolutely can't wait for the updated status message then the solution I gave was the only solution.
If the price of the revocation (which my issuer doesn't charge) is too high for your business then your only option is to wait.
The overall tone of the post seems rather out of proportions though.
Instead, a group of dedicated professionals updated a world-wide infrastructure in less than two days. If you were running your own systems would you have managed that? Yes, you would have known exactly when you were done but could you predict ahead of time when you'd be done?
So, as engineers we make trade-offs and AWS is a pretty clear winner when you look at the TCO of having a scalable architecture. Once you've made that trade-off, the down-side is that you don't have the ultimate flexibility provided by a bare-metal host.
- They managed to create and test a deployment procedure in a couple of hours. - Deployed this update to thousands of machines spread into multiple continents in 48 hours. - There were no downtimes. No action required from customers. - Eveything seems to in in order now.
Yeah.. I hope they die in a fire.
I think the response from amazon was fine and they are clearly not amateurs technically, communication was very poor.
Now that the technical issue has been resolved by other smart people, I'm gonna replace our cert keys and be just fine.
And they actively refused to let people know when they were safe to do that. That is the problem.
http://blog.opbeat.com/posts/amateur-hour-at-aws/few-hours.p...
I want to be clear here that there are plenty of arguments to be made here, but you are not making them.
"No action required from customers. - Eveything seems to in in order now."
What? What about their customers revoking and replacing certs?
It's about the expectations you convey to your customers. If you don't want "amateur" to sting, don't pretend not to be one. Make sure your customers have an accurate understanding of the services you provide and your ability to provide them.
Imagine tomorrow that someone finds a remotely exploitable kernel issue, perhaps involving UDP packet handling. If you have the right infrastructure in place, you should be able to drop the patch file in the right directory and run a script that builds a new system package, runs some automated testing, and then pushes that package out immediately using whatever rolling update strategy is normally used, but at an accelerated pace.
I wish I had time to build something that makes patching system packages on debian systems simpler, making it trivial for businesses to "fork" the distribution as necessary to work around issues (whether they be security critical or not). I've written more thoughts on the matter on my blog: http://stevenjewel.com/2013/10/hacking-open-source/
If you're managing a smaller set of servers, I've been pretty happy with apticron and nullmailer as a way to make sure security updates are applied everywhere. It'd be nice if it could receive notification of security issues faster, perhaps via some sort of push mechanism, but it at least gets things taken care of within 24 hours.
From the Cloudflare blog: "This bug fix is a successful example of what is called responsible disclosure. Instead of disclosing the vulnerability to the public right away, the people notified of the problem tracked down the appropriate stakeholders and gave them a chance to fix the vulnerability before it went public."
-- http://blog.cloudflare.com/staying-ahead-of-openssl-vulnerab...
In this case Amazon choose to focus on fixing the problem versus communicating every detail. The potential consequence in dollars of not fixing the issue in a timely manner is likely in triple-digit millions. I would imagine the cost of having some small number of users complain about how they weren't up to date with communications isn't worth them focusing on it.
Also, I would imagine that customers such as Netflix and other major clients likely got more in-depth communication.
I'm sure it's not going to really affect Amazon's bottom line if opbeat decides to move to a different provider. If you are a very small fish in a big ocean, expect to be treated that way. It's sad to say that but that's the reality.
Personally, I'll take them getting this fixed as a fast as possible, versus getting hourly updates telling me how they're still working it. For example, I just want to hear that they found the MH370 plane. I'm tired of reading news stories with the same depressing message.
Use of Heroku as an example of communications leadership is misplaced. I have had support request sit around unanswered for days in Heroku. Don't get me wrong - I love Heroku. But you have to pay them an arm and a leg monthly in order to get quick support response time. AWS, on the other hand, seems very responsive to all requests.
It's frustrating to wait for updated information, but AWS delivered reasonable details as they were available. Could it improve? Probably. Does it rate worse than other vendors? No, not at all. Consider recent Rackspace or Azure outages, information follows hours later, sometimes days.