There are process-fixes for this, such as requiring a two-person rule when at a production shell and modifying tooling to detect potentially unintentional commands (e.g. a SQL UPDATE without a WHERE) - but given what I know about Amazon's internal practices (i.e. the brutality) it wouldn't surprise me if they did terminate the unfortunate operator - not because they want to, but because AWS simply has too many large-scale customers who would demand immediate action like that.
Everyone who's had operations experience knows that there will be, as time approaches infinity, more than zero SNAFU. That's why companies offer five nines of uptime, not 100% uptime.
1. They were authorized to be doing it. 2. They were following an established process, and not winging it.
The reason for that wording is to illustrate that operational changes are made in accordance with established CM.
Mistakes happen. The system didn't catch the error. That's why the mitigation is "fix the system", and not "fire the team member". =)
Every mistake was used as a learning opportunity to ensure that the same and similar mistakes can't be repeated.
Fixed that by putting the DC domain in red on the prompt.
On a more serious note, if you've never done something like this, you haven't had enough interesting projects.
I've had a decent career and I still managed to:
* re-deploy the current application version in all our data centers, instead of the new version, in a period when our deployment wasn't a 0-downtime one
* rename all the Jenkins jobs on the server to the same name, thus deleting hundreds of Jenkins jobs in one fell swoop
"Let him who is without sin cast the first stone" and all that :)
1. No organization anywhere is a paragon of excellence, and everyone can benefit from improvement. 2. Every organization is made up of humans just like you. With all that entails.
Some things which seem blatantly obvious after the fact are easily overlooked when the pressure to deliver is high and other issues are taking precedence.
I completely agree with your statement. In fact, when I do interviews, one of my favorite and most insightful questions to ask is, basically, "tell me about a time you screwed the pooch." If they don't have a story and they worked in ops, then it can suggest they didn't really do much. The really sharp ones I've interviewed have a good story or two (and can tell it in excruciating detail. =)
* At a prior company I once tried appending to the list of NFS exports, but dropped the "no-root-squash" option, and instantly denied write permissions to our entire VMware farm. You can imagine what then happened to all of the VMs for this mission critical customer. =P
He took off in his piston engine plane, only to lose power during the climb and was forced to make a crash landing. It turned out the airplane was fueled with jet fuel instead of regular gasoline (the ground crewman mistakenly thought the plane was a turbo prop).
Instead of yelling at or firing the ground crewman, Hoover had this to say[2]:
"There isn't a man alive who hasn't made a mistake.
But I'm positive you'll never make this mistake again.
That's why I want to make sure that you're the only one
to refuel my plane tomorrow. I won't let anyone else
on the field touch it."
[1] https://www.aopa.org/news-and-media/all-news/2014/july/pilot... SHUT DOWN S3? ARE YOU SURE? (y/N) :SHUT DOWN S3? ARE YOU SURE? ("I'm absolutely positive."/n) :
"Shut down 73 servers? Are you sure? (y/N)":
"Seventy-three? Wow, I hadn't realized our system grew that much. Probably a new backend dependency got added that I'm not familiar with yet; I'll look into it later." (Y)
Source: I'm was the one the one they would call for our team... usually at 4:00 AM because one of our team members (which was also frequently me) didn't document something correctly.