Summary of the Amazon S3 Service Disruption
aws.amazon.com
aws.amazon.com
It remains amazing to me that even with all the layers of automation, the root cause of most serious deployment problems remain some variant of a fat fingered user.
This is one of the things that happens with windows, getting up a server is so easy, that people believe that they don't have to understand what's under the hood, and then, we get a lot of miss-configuration and operational issues.
EDIT: derp, my bad, I read as "unauthorized" which was "authorized".
" At 9:37AM PST, an authorized S3 team member using an established playbook executed a command which was intended to remove a small number of servers for one of the S3 subsystems that is used by the S3 billing process. Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended."
Anyway, to me this firstly sounds like a "tool" or command that was too powerful with not enough safeguards. Who knows, the command might even be ambiguous.
"Fire you? I just spent $10 million training you!"
typos happen, if the system didn't stop them that's a design issue or accepted risk.
I may have to upgrade that to take the mighty power of Cloud (TM) into account, though. Billions and trillions of fuck ups per second are now well within reach!
If you made that up, I tip my hat off to you as payment for all my future uses of the phrase.
This quote (paraphrased) actually dates all the way back to 1969:
> To err is human; to really foul things up requires a computer.
If the computer knows exactly what actions would be a mistake, why can't it just do the correct actions (those that aren't a mistake) automatically?
That's good UI design in tools with powerful destructive capabilities. You make the UI to do lots of things v.s. the few things you do routinely different enough that there's no mistaking them.
Which is why I'm pointing out that to design UIs like these you should fall back on slightly different UIs depending on the severity of the operation.
Imagine I just entered a command to remove too many servers that will cause an outage:
"Finished removing servers"
(better than no message, I suppose)
vs "Finished removing 8 servers"
(better, it's still too late to prevent my mistake
but at least I can figure out the scale of my mistake)
vs "8 servers will be removed. Press `y` to continue"
(better, no indication of impact but if I'm paying
attention I might catch the mistake)
vs "40% capacity (8 servers) will be removed.
Load will increase by 66% on the remaining 12 servers.
This is above the safety threshold of a 20% increase.
You can override by entering `live dangerously`."
(preemptive safety check--imagine the text is also red so it stands out) "1382345166 agents will be affected. Proceed? (y/n)"
Was that ~100k or ~1M agents? I can't tell unless I count the number of digits, which itself is slow and error-prone. It's worse if I'm in the middle of some high-pressure operation, because this verification detour will break my concentration and maybe I'll forget some important detail.Now if the number is formatted for a human to consume, I don't have to break flow and am much less likely to make an "order-of-magnitude error":
"1,382,345,166 (1.4M) agents will be affected. Proceed? (y/n)"
I always attempt to build tooling & automation and use it during a project, rather than running lots of one-off commands. I find this usually saves me & my team a lot of time over the course of a project, and helps reduce the number of magical incantations I need to keep stored in my limited mental rolodex. I seem to have better outcomes than when I build automation as an afterthought.You could feedback a clarification, but if that happens too often nobody will double check it after they have seen it over and over.
A confirm step in something as sensitive as this operation is important. It won't stop all user error, but it gives a user about to accidentally turn off the lights on US-EAST-1 an opportunity to realize that's what their command will do.
no, because developers still produce code with bugs.
> The dosage you ordered is an order of magnitude greater than the dosage most commonly ordered for this medicine. Continue? y/n
That said, the only way to completely prevent mistakes is to make the tool unable to do anything at all.
(Or to encode every possible meaning of the word "mistake" in your software. If you could do that, you would probably get a Nobel prize for it.)
For securing very dangerous commands, I'd recommend asking the user to retype a phrase composed of random words, or maybe a random 8-character hexadecimal number - something that's different every time, so can't be memorized.
I've almost deleted my heroku production server even though you need to type (or copy paste....ahem...) the full server name (e.g. thawing-temple-23345).
I think the reason was that because in my mind I was 100% sure this was the right server, when the confirmation came up I didn't stopped to look if indeed this was the correct one so I mechanically started to type the name of the server and just a second before I clicked ok, I had this genius idea to double check.... Oh boy... My heart dropped to the floor when I realized what was I about to do.
You could say that indeed Heroku's system of avoiding errors worked correctly....
However the confirmation dialog wasn't what made me stop... Instead it was my past-self's experience screaming at me and remembering me that ONE time where I did fucked up a production server years ago (it cost the company a full day of customers' bids... Imagine the shame of calling all the winning bidders and asking them what price did they end up bidding to win....)
My point is, maybe no number of confirmation dialogs however complex they are, will stop mistakes if the operator is fixed on doing X. If you are working in a semi-autopilot mode because you obviously are very smart and careful (ahem..) you will just do whatever the dialog asks you to do without actually thinking what you are doing.
What, then, will make you stop and verify? My only guess is that experience is the only way. I.e. only when you seriously fuck up you learn that no matter how many safety systems or complex confirmation dialogs there are you still need to double and triple check each character you typed, lest you want to go through that bad experience again....
But I still believe auto-pilot mode is a real thing (and a danger!) .
My point is that I'm not sure if it's even possible to design one that actually cuts errors to 0.
And if that's indeed the case, even if it's close to 0, it's still non-zero, thus at the scale Amazon operates at, it's very probable that it will happen at least one time.
Maybe sometime in the future AI systems will help here?
I've also built complex systems that have been run in production for years with relatively few typo-related problems. The way I do it is with the design patterns like the one I just mentioned, which is also what TeMPOraL was talking about (and I guess you missed it.)
If you have the same kind of confirmation whenever you delete a thing, whether it's an important thing or not, you're designing a system which encourages bad auto-pilot habits.
You'll also note that Amazon's description of the way that they plan on changing their system is intended to fire extra confirmation only when it looks like the operator is about to make a massive mistake. That follows the design pattern I'm suggesting.
Personally, I don't believe it is without making the tool impotent. But you can try and push down the error probability down to arbitrarily low value.
You could go further and try to prevent cat-on-the-keyboard mistakes, which is maybe what you're describing (solve this math equation to prove you are a human who is sufficiently not inebriated). Or even further and prevent malicious, trench-coat wearing, pointy-nosed trouble-makers.
The point is, yes, it is possible. That's what good design does.
One thing I have been doing for my own command line tools is making a preview feature for what a command will do and make the preview state be default. It's simple, but if the S3 engineer first saw a readout of the huge list of servers that were going to be taken offline instead of the small expected list we probably would not be talking about this. There's obviously a ton more you can do here (have the tool throw up "are you sure" messages for unusual inputs, etc).
Simple example: I have a git hook which complains at me if I push to master. If I decide "screw you, I want to push to master", it can't assess my decision, but it easily fixes "oops, I thought I was on my branch".
So, this means a) strong superhuman AI (good luck), b) deciding from an ambiguous input to one of possibly mistaken actions (good luck mapping all possible correct states), or c) drool-proof interface ("It looks you're trying to shut down S3, would you like some help with that?").
TL;DR: yes, but it's a cure worse than the disease.
I don't know. I was suggesting it wasn't realistic to do that, and therefore it wasn't realistic to implement a UI that prevents you making mistakes.
The tool as a whole should incorporate a model of S3. Any action you take through the UI should first be applied to this model, and then the resulting impact analyzed. If the impact is "service goes down", then don't apply the action without raising red flags.
Where I work we use PCS for high availability, and it bugs the heck out of me that a fat-fingered command can bring down a service. PCS knows what the effect of any given command will be, but there's no way (that I know of) to do a "dry run" to see whether your services would remain up afterward.
In practice, it would likely be very hard to make a model of your infrastructure to test against, but I can imagine a tool that would run each query against a set of heuristics, and if any flags pop up, it would make you jump through some hoops to confirm. Such a tool should NEVER have an option to silently confirm, and the only way to adjust a heuristic if it becomes invalid should be formally getting someone from an appropriate department to change it and sign off on it.
By the way, this is how companies acquire red tape. It's like scar tissue.
For instance: https://thenextweb.com/shareables/2014/05/16/emory-universit...
Note: Automation is great, you just can't be sloppy with it. EVER.
edit:fix minor typo
A playbook actually represents a lack of automation for a particular task.
The playbook itself should be automated, with automated tests that validate its correctness.
So, fat-fingering something is imminently possible.
I used to follow runbooks/playbooks written on the internal wiki when I worked at Amazon.
At $work, certain types of frequently-occurring alerts have playbooks that document how the alert in question can be diagnosed and how known causes can be remedied. Something like "Look at Grafana dashboard X. If metric Y is doing this and that thing, the cause is Z. Log on to box 16 and systemctl restart the foo.service."
Too often people will put up with the, "well, we only do this once a month so it's not worth automating". Literally, I script everything now, just in simple bash... if I type a command, I stick it into a script, and then run the script. Over time you go back and modify said script to be better, eventually this turns into more substantive application. At a certain point, around the time that you have more than one loop, are trying to do things based on different error scenarios, it's probably time to turn to rewriting it in another language.
The simplest thing this does for me, is guarantee that all the parameters needed are valid and present before continuing.
Though, an alternative to switching to another language is using xargs well. Writing bash with some immutably has been pretty invaluable for my workflows lately. For example
seq 1 10 | xargs -P10 -I{} ssh $host-{} hostnameMy worst DELETE fail however was:
DELETE * FROM table WHERE [long condition that resolves to true for all records]
Now i write
SELECT or SELECT COUNT(*) over and over again until i see the data i expect and then change it to a DELETE/UPDATE.It's not my personal habit but some folks I know turn off auto commit and BEGIN a transaction every time they enter an interactive SQL sessions. They then default to ROLLBACK at least once before COMMITing them.
That and having a user with read-only permissions or a read replica
root@baz # shutdown now
W: molly-guard: SSH session detected!
Please type in hostname of the machine to shutdown: foo
Good thing I asked; I won't shutdown baz ...
Surprising to see such a simple protection neglected.The alternative is that I use the insurance to pay my legal fees when you sue me for not meeting my uptime guarantees.
It says that their ideal-case failure rate is 11 nines; that's how much you should lose to known, lasting issues like machines failing and cutting over.
Amazon's actual SLA offers 2 nines and 3 nines as the credit thresholds. So they're stating the reliability of their known system, and the rest is for events like this.
For example on GCS (Google's S3)...A storage class specifies how many locations the data is made available. All storage classes share the same durability (chance of google loosing your data) of 99.999999999%, but have different availability (chance of being able to retrieve data).
git commit -m 'typo'
In other words, there's a lack of humility about "unknown unknowns".
The follow-up doesn't bullshit with "extra training to make sure no one does this again", it says (effectively) "we're going to make this impossible to happen again, even if someone makes a mistake".
Is mere extra training the right solution here?
Maybe they need something like the procedure that's used in missile silos:
Not allowing the shutdown system to function at all without the explicit authorization of least two people.
That's a lot more than just extra training, and a lot better than a two-key system.
Probably a bad example. The system was a pain in the ass, so they went and circumvented some of its restrictions.
http://gizmodo.com/for-20-years-the-nuclear-launch-code-at-u...
> Those in the U.S. that had been fitted with the devices, such as ones in the Minuteman Silos, were installed under the close scrutiny of Robert McNamara, JFK's Secretary of Defence. However, The Strategic Air Command greatly resented McNamara's presence and almost as soon as he left, the code to launch the missile's, all 50 of them, was set to 00000000.
> Oh, and in case you actually did forget the code, it was handily written down on a checklist handed out to the soldiers.
Plus, managing humans in a 'rat out' system would be incredibly inefficient. Now you need lots of employees just to listen to the ratting!
Tests and configuration scripts don't prevent all breakage. But when you have them, you can say, "We missed that, let's add it," or "That failed, but it's a false positive. Let's add this edge case to this test."
If you have no automation, tests or auditing systems around running deployments, you can't do any of this.
By the way - this is not just Amazon's problem now. We know the internet has a single point of failure. So does a lot of IoT.
When will we experience the first Suicide DevOps?
https://www.wired.com/2010/11/1110mars-climate-observer-repo...
(Specifically https://www.youtube.com/watch?v=6OalIW1yL-k#t=3m but it's worth watching the whole clip (or even the whole movie) if you haven't seen it before. It's from Terry Gilliam's "Brazil".)
It has? I have yet to see the day where I can neither reach my email provider nor Google nor Hackernews. My local provider might screw up occasionally, or some number of of websites go unreachable for whatever reason. But I fail to come up with anything short of cutting multiple see cables that causes more than 50% of servers to be unreachable to more than 50% of users.
[1]: https://en.wikipedia.org/wiki/Space_Shuttle_Challenger_disas...
Amazon is taking the right approach here. The fact that a system as complex and important as S3 can be taken down is a failure of the system, not the person who took it down accidentally.
1. https://en.wikipedia.org/wiki/Capability_Maturity_Model#Leve...
The certification is more for the organization/unit and the people working do not realize what they are for. Another thing that usually becomes a problem is the rigidity of the certification. Saying you need X, Y and Z documented is easy, but it doesn't work for projects that maybe don't have Y. So people make up documentation and process just to be compliant, this soon becomes a hinderance to the work. At this point people either abandon the process or follow it and the work suffers.
(I lied about the "insta" part)
Look at the recent GitLab incident - one guy messed up and nuked a server. Okay, that happens sometimes, go to backups. Uh oh, all the backups are broken. Minor momentary problem just turned into a major multi-day one.
That's a problem and one which could be preventable with training (or, arguably, firing and hiring). Maintaining your backups properly should be someone's duty, designing and testing systems to minimize impact of user error should be too.
If someone doesn't test their backups, you train them to test backups. If someone lies about testing the backups, maybe you fire them. But if someone trips and shatters the only backup disk, you don't yell at them - you create backups that an instant of clumsiness can't ruin.
I did overstate, training is perfectly reasonable, but I often see it cited exactly when it shouldn't be, as a solution to errors like typos or forgetfulness.
Instead, you make a machine verify the backups simply by using the backups all the time. For example, at work I feed part of our data pipeline with backups: Those processes have no access to the live data. If the backups break, those processes would provide bad information to the users, and people would come complaining in a matter of minutes.
Just like when you have a set of backup servers, you don't leave them collecting dust, or tell someone to go look at them every once in a while: you just route 1% of the traffic through them. They are still extra capacity, you can still do all kinds of things to them without too much trouble, but you know they are always working.
Never, ever, force people to do things they don't gain anything from. Their discipline would fade, just like it fades when you force them to a project management tool they get no value from.
One would only actually test the backups about twice a year just to be damn sure they are still resulting in restorable data. The rest of the year it's only worth keeping an automated process reporting whether or not the things are being made, and people keeping an eye on change management to be sure no changes are made to the known-to-be working process that can break it without the new process incurring an explicit vetting cycle. Gitlab wasn't apparently testing or engaging in monitoring what was supposed to be an automated process. That's where they got burned.
Process monitoring may be boring as hell, but it's seldom wasted effort, and will prevent massive, compounded headaches from bringing operations to a chaotic halt.
Nope. Nope. Nope.
You test every backup by automatically restoring from it in a sandbox and verifying its integrity and functionality in the restored state.
Backups are worthless unless verified for their intended use of recovering a functioning system.
And constant "this succeeded" messages don't scale well.
Someone rm -rf / ing the server will happen eventually with near 100% certainty in any company and can be mitigated by tested, regular, multiply redundant backups.
Cosmic rays flipping bits will happen with near 100% probability at the scale someone like Amazon works at can by mitigated by redundant copies and filesystems with checksum style checks. Similar with hard drive failure.
Earthquakes will happen in some areas with near certainty over the time periods companies like Amazon presumably hope to be in business and could be mitigated by having multiple datacenters and well constructed buildings. Similar for 'normal' scale volcanoes.
Fires will happen but they can be mitigated (with appropriate buildings and redundancy).
Small meteorite stikes are unlikely but can be mitigated by redundancy.
Solar activity causing an electomagnetic storm - yeah one can shield one's datacenter in a Faraday cage but in this situation the whole world is probably in chaos and one's datacenter will be the least of one's concerns (unless shielding become standard in which case you'd better be doing it). Similar applies for nuclear war, super volcanoes, massive meteorite strikes or other global events at the interesting end of the scale.
But yeah there are going to be things that get missed. They key is having an organization that (1) learns from its mistakes and (2) learns from others' mistakes and continually keeps their risk modeling and mitigation measures up to date. And note that many of the hazards that are worth mitigating have the same mitigation i.e. redundancy (at different scales).
That's a great line. How should I attribute it?
He won't make the same mistake because no one makes the same big mistake twice? I wouldn't bank on that alone.
Based on Amazon's decision to improve the tooling such that this category of error would be (hopefully) impossible to reproduce, I would lean more towards that being the case.
Enter Frans Plugge. Whenever a customer would get into that mode we'd fire Frans. This was easy, simply because he didn't exist in the first place (his name was pulled from a skit by two Dutch comedians, bonus points if you know who and which skit).
This usually then caused the customer to recant on how he/she never meant for anybody to get fired...
It was a funny solution and we got away with it for years, for one because it was pretty rare to get customers that mad to begin with and for another because Frans never wrote any blog posts about it ;)
But I was always waiting for that call from the labor board to ask why we fired someone for who there was no record of employment.
It irks me that businesses fire people because of pressure from clients or social media. But having never been the boss, I may be missing something.
Internal repercussions notwithstanding, externally the company is a united front. It cannot cause mistakes by luck, accident, or happenstance, because the world includes luck, accidents, and happenstance, so any user-visible error is ipso facto a failure of management.
It's still mind blowing and very amusing that this is a thing in our world!
Cause that sounds pretty great.
The problem of user error can be mitigated by an appropriate level of OCD.
But OCD can't be trained, you either have it or you don't.
Jeff Bezos once said: "Good intentions never work, you need good mechanisms to make anything happen"
I've had the privilege of either working for myself, the company that acquired mine and let me run the dev, or at Google. From that perspective, and what I understand about ops, the rarity is not having the attitude mentioned in the parent.
I seem to recall an EC2 or S3 outage a few years ago that boiled down to an engineer pushing out a patch that broke an entire region when it was supposed to be a phased deployment.
I could be mis-remembering that but it's important that these lessons be applied across the whole company (at least AWS) so it would be a bigger mark against AWS if this is a result of similar tooling to what caused a previous outage.
(Source: am a self-identified post-mortems connoisseur. :)
You have some process that starts out being "deploy this app with this java code". You deploy once and while, so it's not a big deal. But then those changes get a bit more frequent and so you pull out the common bits and the process becomes "make this YAML change in git and redeploy the app".
That works until you find yourself deploying 5 times a day, so you turn it into a MySQL table, and the process becomes "write a ROLL plan that executes this UPDATE x=y WHERE t=u; command"
After a while you get super annoyed at some quirk of the commands and figure, "Ok, fine, I'll just add an endpoint and some logic that just does this for the command case."
Then you wanna go on vacation and the new guy messed up the API request last week, so you figure, "I'll just add a little JS interface with a little red warning if the request is messed up in this way or that before I go".
You get back from vacation and some original interested party (whoever has wanted all these changes deployed) watched the intern make the change and thinks they could just do it themselves if they had access to the interface. You're wary, but you make the changes together a few times and maybe even add a little "wait-for-approval" node in the state machine.
Life is good. You've basically de-looped yourself, aside from a quick sanity check and button press, instead of what was a ~2 hour code + build + PR + PR approved + deploy process.
Then that interested party goes to work for Uber and the rest of your team adds a few functionalities on top of the interface you built and it all goes pretty well, until you realize that now that this thing that used to be 20 YAML objects is now 50k database records, and a bunch of them don't even apply anymore. So you build a button to disable some group of them, but after getting it deployed you realize it's actually possible to issue a "disable all" request accidentally if you click a button in your janky JS front-end before the 50k records download and get parsed and displayed. Oops! This mistake that you and the original interested party would have never made (because you spent the last 2 years thinking about all this crap) is probably a single impatient anxious mouse-click away from happening. So you make a patch and deploy that.
Congrats! You found that particular failure mode and added some protections for it, and maybe added some other protections like rate-limiting the deletions or updates or whatever. That's cool, but is that every failure mode? I bet it isn't. What happens when someone else thinks you have too many endpoints and just drops to SQL for the update?
Basically, yeah, of course you think of this stuff while iterating on it. But you figure "only power users are on the ACL" or "my teammates will understand the data model before making changes, or ask me first" or "that's what ROLL plans are for" or "I'll show a warning in the UI" or whatever. Fundamentally, you're thinking about a way to do a thing, if you're even thinking about it at all.
So yeah, that's what I've spent the last year or two doing. :-)
Those lines are not reflective of what Amazon is but what picture Amazon wants to paint now. They have clarified it is their error and not some hacking attempt. Secondly they have not vilified the engineer in question because already Amazon's culture is a bit of a ??? in public mind.
But they have got it right. Shit happens and this is not the first time it has happened or the last time it will happen. Also it will happen with Microsoft, Google and everyone else.
May be we will build even better technologies that will rely on two different cloud providers instead of 1.
Actually the accounts I've read seem to indicate that most missile operators simply decided they would never launch no matter what. God bless them, for that.
This one did not carry a warhead. Others do...
Turns out the answer to your question is simply: luck.
I think Borland do some RAD systems, and Microsoft have an IDE of sorts on the way too.
EDIT: Please note that this is humour.
I suspect locking everyone down in the way you suggest would cost more in lost productivity (and costs for the infrastructure that would be required for greater auditing, etc.) than is lost in outages like this.
> While this is an operation that we have relied on to maintain our systems since the launch of S3, we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years.
These sorts of things make me understand why the Netflix "Chaos Gorilla" style of operating is so important. As they say in this post:
> We build our systems with the assumption that things will occasionally fail
Failure at every level has to be simulated pretty often to understand how to handle it, and it is a really difficult problem to solve well.
http://techblog.netflix.com/2011/07/netflix-simian-army.html
> Chaos Gorilla is similar to Chaos Monkey, but simulates an outage of an entire Amazon availability zone. We want to verify that our services automatically re-balance to the functional availability zones without user-visible impact or manual intervention.
Exactly. It seems likely that Amazon tests the restart operation, but it would be hard to test it at full us-east-1 scale. Running a full S3 test cluster at that scale would likely be a prohibitive expense. Perhaps the "index subsystem" and "placement subsystem" are small enough for full-scale tests to be tractable, but certainly not cheap, and how often do you run it? Also, hindsight is 20/20, but before this incident it might have been hard to identify "full-scale restart of the index subsystem" as rising to the top of the list of things to test.
One approach is to try to extrapolate from smaller-scale tests. It would be interesting to know what kinds of disaster testing Amazon does do, and at what scale, and whether a careful reading could have predicted this outcome.
Rough guide:
CT = cost of 1 full scale test with necessary infrastructure and labor costs added up
CF = amount of money paid out in SLA claims + subjective estimate of business lost due to reputation damage etc
PF = estimate of probability of this event happening in a given year
if PF * CF > CT, then you run such a test at least once a year. Think of such an expense as an insurance premium.
What Netflix does with their simian army is amortize the cost of doing the test across millions of tests per year and the extra design complications arising from having to deal with failures that often.
(Nit: this incident affected a region, not a zone. us-east-1 is a region, which is divided into zones us-east-1a, us-east-1b, etc. S3 operates on regions.)
Keep in mind, S3 "fails" all the time. We regularly make millions of S3 requests at my work. Usually we get 1:240K failure rate (mostly GETs), returning 500 errors. However, if you're really hammering an S3 node in the hash ring (e.g. Spark job), we see failures in the 1/10K range, including SocketExceptions, where the routed IP is dead.
You need to always expect such services to die in your code, setting the proper timeouts, backoffs, retries, queues, and dead letter queues.
Sometimes it's a 404 for an object written 1 sec prior, other times it's an S3 node that died mid request. Retry gets you to a different node.
Ensuring that your status dashboard doesn't depend on the thing it's monitoring is probably the first thing you should think about when designing your status system. This doesn't fill me with confidence about how the rest of the system is designed, frankly...
Since restarting the entire fleet would incur downtime of all relevant S3 operations, it's unlikely that it was something ever intentionally done in production (and they may or may not have run that scenario in other environments).
Source: I used to run several large scale services at Amazon.
What makes you think they didn't?
The website that shows the public results of the monitoring, which is updates only by humans, depended on it.
My us-east-1 RSS feed said S3 had no incidents.
"we have changed the SHD administration console to run across multiple AWS regions."
That being said, a better solution would be to stand up a separate infrastructure just for status pages.
What they really need is failover capability, which can fire up the status page on a competitor's service (or maybe on a completely separate disaster recovery site site owned by Amazon) in case Amazon's own services go down.
I'm sure Amazon's architects and engineers are more than capable of designing and implementing such a robust system and recognizing its importance. So it puzzles me as to why it wasn't done.
That was CEO Robert Allen's response when the AT&T network collapsed [1] on January 15, 1990
He was asked who made the mistake.
I can't imagine any CEO now a days making a similar statement.
[1] http://users.csc.calpoly.edu/~jdalbey/SWE/Papers/att_collaps...
We all watched the news and I recall him saying that. The specific quote I don't remember but it was something like " you can consider that I did." I think he was asked what will happen to the person that caused it and who is that person.
Everyone knew right away this had to be human error. Right away. Switches simply had too much redundancy.
It was big then and not sure if I can locate a video.
> As far as our customers are concerned, I did it.
http://www.upi.com/Archives/1990/01/16/ATT-pinpoints-cause-o...
I admired him and his answer at the time. The culture was quite professional and blame really never existed.
[0]: http://www.cnbc.com/2009/04/30/Portfolios-Worst-American-CEO...
both better for morale, and better for preventing another incident.
I'd be far more impressed if a low-level employee who's whole family depended on his job and who stood a good chance of getting fired admitted a serious mistake.
To me that is just another example of 'caring theater'. Whereby carefully crafted PR responses [1] appear to take responsibility in a 'buck stops here' kind of way. The truth is it is unreasonable in many cases for the top person to be able to prevent any and all errors. If you try and make everything perfect with no mistakes you would never make any money (and of course it's not even possible).
[1] ie 'our customers safety and security is of the utmost importance to us'.
It sounds like the weakness in the process is that the tool they were using permitted destructive operations like that. The passage that stuck out to me: "in this instance, the tool used allowed too much capacity to be removed too quickly. We have modified this tool to remove capacity more slowly and added safeguards to prevent capacity from being removed when it will take any subsystem below its minimum required capacity level."
At the organizational level, I guess it wasn't rated as all that likely that someone would try to remove capacity that would take a subsystem below its minimum. Building in a safeguard now makes sense as this new data point probably indicates that the likelihood of accidental deletion is higher than they had estimated.
s3-shutdown -c "one hundred fifty"
But something simpler like a --emergency flag or the more whimsical --shutitdownshutitalldown[0] https://www.quora.com/When-people-who-work-at-Facebook-say-c...
I've always wondered why ops hasn't adopted some of the best practices that have been around for years to avoid fat finger errrors. Like why don't we have systems where to do something dangerous requires two separate people run the command, or there's an approval step, or whatever.
I've always found that explanation a little threadbare.
There are two distinct failures here, the tool being too liberal and then the failure mode not being well understood for this index subsystem.
If I remember correctly, all of our processes were organized through a weekly change management process that went through reviews of the exact commands to be run. Being oncall was a bit more liberal, you would typically execute commands on production as needed based on your experience and with others over your shoulder if you had any doubt. Interacting with the gossip protocol was a pretty common thing to do when you were triaging issues.
Unrelated, I was briefly a SME on the EBS billing system and probably interacted with the poor guy whom executed this command.
Something similar happened to us, when an Engineer deleted part of our production database with a single command. Fortunately, we could reconstruct it from backups and replication logs.
There is no easily readable timeline. It is not discoverable from anywhere outside of social media or directly searching for it. As far as I know, customers were not emailed about this - I certainly wasn't.
You're an important business, AWS. Burying outage retrospectives and live service health data is what I expect from a much smaller shop, not the leader in cloud computing. We should all demand better.
AWS has more implicit trust that this won't happen again, since they've never (I think?) had something like this happen, so just a few lines about fixing the tool that let all the nodes shutdown is enough to restore confidence.
A graphical illustration of the service dependencies they were talking about would have been nice as well.
To receive a Service Credit, you must submit a claim by opening a case in the AWS Support Center. To be eligible, the credit request must be received by us by the end of the second billing cycle after which the incident occurred and must include:
the words “SLA Credit Request” in the subject line; the dates and times of each incident of non-zero Error Rates that you are claiming; and your request logs that document the errors and corroborate your claimed outage (any confidential or sensitive information in these logs should be removed or replaced with asterisks). If the Monthly Uptime Percentage applicable to the month of such request is confirmed by us and is less than the applicable Service Commitment, then we will issue the Service Credit to you within one billing cycle following the month in which your request is confirmed by us. Your failure to provide the request and other information as required above will disqualify you from receiving a Service Credit."
A link to their tweet about the status page not working because the building was burning down around it seems compelling.
I find making errors on production when you think you're on staging are a big one for similar errors. One of the best things I ever did on one job was to change the deployment script so that when you deployed you would get a prompt saying "Are you sure you want to deploy to production? Type 'production' to confirm". This helped stop several "oh my god, no!" situations when you repeated previous commands without thinking. For cases where you need to use SSH as well (best avoided but not always practical), it helps to use different colours, login banners and prompts for the terminals.
That's what I meant about the SSH comment. Not every team has the automation or infrastructure that allows you to avoid SSH.
If the prompt was type "production" to confirm, I'm sure I'd just as readily train myself to jump the gun on that one.
This is analogous to "we needed to fsck, and nobody realized how long that would take".
I felt TERRIBLE about it.
I did all this while sitting 2 feet from a print out of "The 8 fallacies of distributed systems". Bandwidth is indeed not infinite, can confirm.
Usual checks to access $HOSTNAME failed
Rushed to office at 6am before important process was about to run that needed that host.
Plugged in keyboard+monitor, dead screen, nothing.
Physically power-cycled server.
Stood in front of monitor+keyboard. It occurred to me it was taking longer than expected to show POST screen. About that time, I got a page saying $ACTUALHOSTNAME is down.
Walk around to the back of the racks. The monitor cable had come detached from the cable extender that I plugged into the server. I had never plugged the monitor in at all, just the extension.
The server wasn't down in the first place, it just lost a virtual interface, which I was paged for, and stupidly tested that virtual interface instead of the REAL name/IP.
And then I raced to the office just so that I could cause an outage.
The only thing I read in there and go "hmmm" is that it took quite that long for the S3 service to recover, and that the status page wasn't hosted on someone that doesn't have an S3 dependency. That's just a plain "doh" moment :)
Dear Amazon: please lease a $25/month dedicated server to host your status page on.
The lesson is that partitioning your service into isolated regions is not enough. You need to partition your load evenly, too. I can think of several ways to accomplish this:
1. Adjust pricing to incentivize customers to move load away from overloaded regions. Amazon has historically done the opposite of this by offering cheaper prices in us-east-1.
2. Calculate a good default region for each customer and show that in all documentation, in the AWS console, and in code examples.
3. Provide tools to help customers choose the right region for their service. Example: http://www.cloudping.info/ (shameless plug).
4. Split the large regions into isolated partitions and allocate customers evenly across them. For example, split us-east-1 into 10 different isolated partitions. Each customer is assigned to a particular partition when they create their account. When they use services, they will use the instances of the services from their assigned partition.
> Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended.
If I would have guessed anyone could prevent mistakes like this from propagating it would be AWS. It points to just how easy it is to make these errors. I am sure that the SRE who made this mistake is amazing and competent and just had one bad moment.
While I hope that AWS would be as understanding as Gitlab, I doubt the outcome is the same.
I hadn't thought about how 'prone to forgive' an organization is.
Should there be policies in place as to how big a mistake you are allowed to commit? Something akin to:
```if(log10(cost_of_mistake) >= 5): worker.fire() ```
Source: AWS employee
Self-hosting your diagnostic tools is an easy mistake to make, and I've seen both startups and large, multi-decade-experienced companies make it.
It'd be very interesting to know what kind of tech they use at AWS to throttle or do circuit breaking to allow back-end services like the indexer to come up in a manageable way.
It would be easy for an arrogant organisation to fire or negatively impact the person that made the mistake, I hope Amazon don't fall into that trap and focus instead on learning from what happened, closing the book and move on.
I wouldn't task a junior sysadmin a server deletion, would you? Nor could I ever consider someone without a fuckup a senior ;)
I'm genuinely curious. As my experiments with it have left me disappointed with its performance, I'm just not sure what I could use it for. Store massive amounts of data that is infrequently accessed? Well, unfortunately the upload speed I got to the standard rating one was so abysmal it would take too much time to move the data there; and then I suspect the inverse would be pretty bad as well.
https://aws.amazon.com/blogs/aws/aws-storage-update-amazon-s...
If you have lots of data that needs to be uploads (TB/PB worth), then I'd take a look at AWS Snowball. https://aws.amazon.com/snowball/
Also, if you're using the AWS CLI to upload, make sure multi-part upload is enabled.
Why use S3? It scales without any user interaction, is highly available (yes, I cringe saying that, but this is the first, and hopefully only time this has occurred! =D ) and extremely easy to access; it's as simple as an HTTP GET. Being able to address objects directly and not have to worry about managing file systems simplifies a lot.
All those tweets saying "turn it off and back on again"?
"We accidentally turned it off, but it hasn't been turned it off for so long it took us hours to figure out how to turn it back on."
Poorly-presented jokes aside, this is rather concerning. The indexer and placement systems are SPOFs!! I mean, I'd presume these subsystems had ultra-low-latency hot failover, but this says they never restarted, and I wonder if AWS didn't simply invest a ton of magic pixie dust in making Absolutely Totally Sure™ the subsystems physically, literally never crashed in years. Impressive engineering but also very scary.
At least they've restarted it now.
And I'm guessing the current hires now know a lot about the indexer and placer, which won't do any harm to the sharding effort (I presume this'll be being sharded quicksmart).
I wonder if all the approval guys just photocopied their signatures onto a run of blank forms, heheh.
The system is a collection of shards. If you replicate it to create a second shard, then you'll just have a large a system, which is still a single point of failure.
The index, by necessity, has to be able to answer the question 'this object exists' or 'this object doesn't exit' - so it needs to have consensus.
My speculative presumption was going off the sole datapoint of "we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years". I'm not quite sure how to interpret "restart" in this context, mostly due to lack of exposure or experience.
The report also says "Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended. The servers that were inadvertently removed supported two other S3 subsystems." So you're right, it looks like multiple servers were supporting these systems, which does make sense (especially considering the load they would have seen). Okay.
I guess I didn't quite think through the load requirements and thought these were single machines - which is certainly ludicrous thinking :) - and that's where I got the SPOF reasoning from.
You're very right though, these consensus systems must be built as bottlenecks in order to see everything.
And there aren't really any alternatives: "build extra indexers and placement systems!" just gives you "but what if _all_ of them get taken offline?" and "it can't leave the datacenter, it sees 100GB/s of throughput" (number taken out of thin air).
Good points.
CEO's all over the world just realized that they can't only depend on S3, and they might have to double up on their infrastructure and have a parallel env. on Azure or Google as well.
So while it is perhaps necessary to be multi-regional in Amazon, you wouldn't necessarily have to go multi-provider.
the most important one is that Spanner runs on Google’s private network. Unlike most wide-area networks, and especially the public internet, Google controls the entire network and thus can ensure redundancy of hardware and paths, and can also control upgrades and operations in general. Fibers will still be cut, and equipment will fail, but the overall system remains quite robust. It also took years of operational improvements to get to this point. For much of the last decade, Google has improved its redundancy, its fault containment and, above all, its processes for evolution. We found that the network contributed less than 10% of Spanner’s already rare outages.
But when it fails it's going to be epic!
[1] https://cloudplatform.googleblog.com/2017/02/inside-Cloud-Sp...
From a software development perspective, it makes sense to reuse S3 and rely on it internally if you need object storage, but from an ops perspective, it means that S3 is now a single point of failure and that SES's reliability will always be capped by S3's reliability. From a customer perspective, the hard dependency between SES and S3 is not obvious and is disappointing.
The whole internet was talking about S3 when the AWS status dashboard did not show any outage, but very few people mentioned other services such as SES. Next time we encounter errors with SES, should we check for hints of S3 outage before everything else? Should we also check for EC2 outage?
I don't think this is particularly surprising. I'd already pretty much assumed that, e.g., a package of code for a Lambda function would be housed in an S3 bucket somewhere.
What's really surprising to me is how many of those buckets appear to live in US-EAST-1, and aren't able to keep functioning in a catastrophe by failing over to a different region.
We're in ap-southeast-2 (Sydney) and none of our services were impacted yesterday.
On a more serious note, if you've never done something like this, you haven't had enough interesting projects.
I've had a decent career and I still managed to:
* re-deploy the current application version in all our data centers, instead of the new version, in a period when our deployment wasn't a 0-downtime one
* rename all the Jenkins jobs on the server to the same name, thus deleting hundreds of Jenkins jobs in one fell swoop
"Let him who is without sin cast the first stone" and all that :)
1. No organization anywhere is a paragon of excellence, and everyone can benefit from improvement. 2. Every organization is made up of humans just like you. With all that entails.
Some things which seem blatantly obvious after the fact are easily overlooked when the pressure to deliver is high and other issues are taking precedence.
I completely agree with your statement. In fact, when I do interviews, one of my favorite and most insightful questions to ask is, basically, "tell me about a time you screwed the pooch." If they don't have a story and they worked in ops, then it can suggest they didn't really do much. The really sharp ones I've interviewed have a good story or two (and can tell it in excruciating detail. =)
* At a prior company I once tried appending to the list of NFS exports, but dropped the "no-root-squash" option, and instantly denied write permissions to our entire VMware farm. You can imagine what then happened to all of the VMs for this mission critical customer. =P
He took off in his piston engine plane, only to lose power during the climb and was forced to make a crash landing. It turned out the airplane was fueled with jet fuel instead of regular gasoline (the ground crewman mistakenly thought the plane was a turbo prop).
Instead of yelling at or firing the ground crewman, Hoover had this to say[2]:
"There isn't a man alive who hasn't made a mistake.
But I'm positive you'll never make this mistake again.
That's why I want to make sure that you're the only one
to refuel my plane tomorrow. I won't let anyone else
on the field touch it."
[1] https://www.aopa.org/news-and-media/all-news/2014/july/pilot... SHUT DOWN S3? ARE YOU SURE? (y/N) :SHUT DOWN S3? ARE YOU SURE? ("I'm absolutely positive."/n) :
"Shut down 73 servers? Are you sure? (y/N)":
"Seventy-three? Wow, I hadn't realized our system grew that much. Probably a new backend dependency got added that I'm not familiar with yet; I'll look into it later." (Y)
Source: I'm was the one the one they would call for our team... usually at 4:00 AM because one of our team members (which was also frequently me) didn't document something correctly.
Every mistake was used as a learning opportunity to ensure that the same and similar mistakes can't be repeated.
Fixed that by putting the DC domain in red on the prompt.
There are process-fixes for this, such as requiring a two-person rule when at a production shell and modifying tooling to detect potentially unintentional commands (e.g. a SQL UPDATE without a WHERE) - but given what I know about Amazon's internal practices (i.e. the brutality) it wouldn't surprise me if they did terminate the unfortunate operator - not because they want to, but because AWS simply has too many large-scale customers who would demand immediate action like that.
Everyone who's had operations experience knows that there will be, as time approaches infinity, more than zero SNAFU. That's why companies offer five nines of uptime, not 100% uptime.
1. They were authorized to be doing it. 2. They were following an established process, and not winging it.
The reason for that wording is to illustrate that operational changes are made in accordance with established CM.
Mistakes happen. The system didn't catch the error. That's why the mitigation is "fix the system", and not "fire the team member". =)
Is that code for a "did you try to reboot the system?" kind of troubleshooting?
It sounds to me like the authorized engineer sent a command to reboot/reimage a large swath of the S3 infrastructure.
I would add, it would be awesome if there was a simulation environment, beyond just a test environment that simulated servers outside requesting in, before a command was allowed to run onto production, like a robot deciding this, then could mitigate this, kind of like TDD on steriods if they don't have that already.
How do surgeons react when they accidentally cut the wrong thing?
Marsh chronicles his career, and includes (at least) one story about the slip of the scalpel and the... result.
Highly recommended.
I'll make sure my soon-to-be MD girlfriend reads it too
I was the guy who deployed the update to every one of our servers as I walked out the door for the day. So I know what this feels like.
You learn to get through it fast, because there's no other choice, you're almost always the best placed person to clean up the mess.
And one day you can look back and the scars have healed.
My guts hurt just reading this.
With big failures is never just one thing. There are a series of mistakes, bad choices, and ignorance that lead to a big system wide failures.
We have geo-distributed systems. Load balancing and automatic failover. We agonize over edge cases that might cause issues. We build robust systems.
At the end of he day reliability -- a lot like security -- is most affected by the human factor.
I'd be interested to understand why a cold restart was needed in the first place. That seems like kind of a big deal. I can understand many reasons why it might be necessary, but that seems like one of the issues that's important to address.
In this case, throwing away and then re-provisioning the split-off nodes is a viable approach.
Yeah... nothing says "resilience" quite like that...
It's good practice in general, and I'm kind of astonished it's not part of the operational procedures in AWS, as this would have quickly been caught and fixed before ever going out to production.
There's no way this could have been mitigated with a dry run. They're mitigating it in future by putting more aggressive safeguards in their tooling, which is the correct way to mitigate this sort of issue.
For numerical inputs, one might use both the digits and the textual expression. This would make them quite cumbersone but much less prone to errors. Or devise some shorthand for them...
156 (on fi six). 35. (zer th fi). 170 (on se zer). 28 (two eig) evens have three letters odds have two.
This is just my 2 cents.
just like rm -rf / should really be rm -rf `root`
This is the bit that'd worry me most; you'd think they'd be testing this.
Moments like these always remind me that a particularly clever or nefarious set of individuals could shut down essential parts of the Internet with a few surgical incisions.
Shameless plugs (authored months ago): http://tuxlabs.com/?p=380 - How To: Maximize Availability Effeciently Using AWS Availability Zones ( note read it, its not just about AZ's it is very clear to state multi-regions and better yet multi-cloud segway...second article) http://tuxlabs.com/?p=430 - AWS, Google Cloud, Azure and the singularity of the future Internet
This is a case of someone slipping on the keyboard, removing more capacity than intended and the recovery process taking longer than expected. The process actually seems to be working (to a given value of working), but the amount of downtime was way above acceptable. They've already put more safeguards into the tooling to prevent the situation from happening again.
S3 is also orders of magnitude more complex than Gitlabs infrastructure, so while the amount of time the outage lasted for is not acceptable, it does show that they at least have working processes for critical situations that allow them to get back in service within a day, which is pretty impressive.
In this particular case the scripts didn't have adequate protections in place, but that's the benefit of hindsight
Why is good design only possible in GUIs?
A good command line would better protection would have helped, but this fails some of the core of good interface design.
That is EXACTLY what they are doing (among other things).
https://news.ycombinator.com/newsguidelines.html
We detached this comment from https://news.ycombinator.com/item?id=13776335 and marked it off-topic.
Never type 'EXEC DeleteStuff ALL'
When you actually mean 'EXEC DeleteStuff SOME'
>We understand that the SHD provides important visibility to our customers during operational events and we have changed the SHD administration console to run across multiple AWS regions.
"From the beginning of this event until 11:37AM PST, we were unable to update the individual services’ status on the AWS Service Health Dashboard (SHD) because of a dependency the SHD administration console has on Amazon S3. Instead, we used the AWS Twitter feed (@AWSCloud) and SHD banner text to communicate status until we were able to update the individual services’ status on the SHD. We understand that the SHD provides important visibility to our customers during operational events and we have changed the SHD administration console to run across multiple AWS regions."
Shameless plugs (authored months ago): http://tuxlabs.com/?p=380 - How To: Maximize Availability Effeciently Using AWS Availability Zones ( note read it, its not just about AZ's it is very clear to state multi-regions and better yet multi-cloud segway...second article)
http://tuxlabs.com/?p=430 - AWS, Google Cloud, Azure and the singularity of the future Internet
* "authorized S3 team member" -- how did this team member acquire these elevated privs?
* Running playbooks is done by one member without a second set of eyes or approval?
* "we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years"
The good news:
* "The S3 team had planned further partitioning of the index subsystem later this year. We are reprioritizing that work to begin immediately."
The truly embarrassing that everyone has known about for years is the status page:
* "we were unable to update the individual services’ status on the AWS Service Health Dashboard "
When there is a wildly-popular Chrome plugin to fix your page ("Real AWS Status") you would think a company as responsive as AWS would have fixed this years ago.
This is not in the post.
) If it's a playbook for something with minimal intended impact sure. The issue is that the tooling had larger capabilities than should be.
) Yes that seems like a major, major problem.