Postmortem for outage of us-east-1
joyent.com
joyent.com
"...we will be rethinking what tools are necessary over
the coming days and weeks so that "full power" tools
are not the only means by which to accomplish routine tasks."Everyone makes mistakes, and judging by the language of the postmortem it appears it was just that.
Would love to have that sysadmin on my team, because he will never do that again....
She's probably less likely than before due to fear or guilt, but I don't know that this makes her less error prone than every other person in the country, including those that double check every time, for instance.
There has been no evidence of them doing this at all.. I don't know why people keep saying it.
This bit bothers me more than anything else. It's not just a die roll configuration, but a known one with an operator required to do the re-rolling.
Everything is a value judgement I guess, but knowingly leaving that one 'mitigated' would drive me insane.
Coincidentally we had a visit from the CTO of high-performance storage vendor who happen to have a bunch of kernel hackers on their staff. We mentioned our problem in passing and he explained how they'd nearly lost a major contract because a deployment had moved thei customer from being storage-bound to initially untracable data loss. Digging around by their kernel hackers showed the NIC was losing data. After a certain amount of too-ing-and-fro-ing the vendor moved from denying the problem to admitting that their silicon had a defect that would throw away data under load, and their Windows drivers tried to spackle over the problem. They were relying on no-one being able to drive enough load through the card to cause a problem.
This dovetailed with our experience, and we found that installing a different manufacturer's card in the blade let us work around the problem. Our blade supplier moved to a new NIC vendor subsequently.
"To make error is human. To propagate error to all server in automatic way is #devops." -@DEVOPS_BORAT [1]
[1]https://twitter.com/DEVOPS_BORAT/status/41587168870797312
"Because there was a simultaneous reboot of every system in the datacenter, there was extremely high contention on the TFTP boot infrastructure, which like all of our infrastructure, normally has throttles in place to ensure that it cannot run away with a machine."
What does "cannot run away with a machine" mean? Why would you want to by default restrict the speed at which that system runs?
The reason why is probably related to what it does besides TFTP, which under normal circumstances is probably more important than new nodes joining the network.
During a mass reboot that will bite you because then the chances of starvation will go up quite a bit as the throttles cause one machine after another to time-out and retry their boot sequence.
[1] http://www.amazon.com/The-AWK-Programming-Language-Alfred/dp...
Substitute reboot with 'upgrade to win 7' and data center with university at get a story from a month? ago.
rm ./*.* - delete all files in current dir
rm /*.* - never do this
Can you spot the difference? A colleague did the later on our testing server. No clients were disturbed, but us devs were left working on things that can run locally (much more pleasant stuff for sure, yet the schedule suffered a lot).There are a large number of varieties of this particular error. Some with terrible results.
rm -rf * .bak
for instance (especially when executed in the root directory).
That '#' prompt is there for a reason.
The way to solve these sort of issues is to first get the files using 'find' until you're totally happy about the result and then to use 'rm' as the command passed to find.
Of course, nobody does this ;)
Would also erase everything else on the box so you might not know for a few seconds, but then the database errors start and the lib and pid files start disappearing and now your the king of a mountain of shit. It won't take till next reboot to notice.
If you ever get a chance, rm -rf / a box before you throw it out and just play around with it for a couple minutes while it eats itself.
You can spare your self most errors like that with tab complete, the built in sanity checker. I tab complete everything since I'm dyslexic, and it saves me a shit ton of time cause I'm never far from the error when I notice it.
Yes, it will, because the original post didn't include -r. So it's only deleting things that match the glob in the root directory. On many systems, that is nothing.
/*.*
only matches files and directories in / that have a period in their name. So it would skip over /usr, /home, etc.Try
chown -R user:group /var/local/some/data/dir /
inside an init script. Of course it had been tested many times, and the space was only fat-fingered in when someone move the data directory.Instead of /usr/some/path/here
http://www.miltonbayer.com/blog/news/when-a-code-commit-goes...
A colleague of mine once called me over because they were having trouble with their computer. Apparently while doing "a little tidying up" they decided to move the windows folder to somewhere else. I can't remember the details but for some reason it was impossible to get a command prompt (Windows NT maybe?) without a disk, which I didn't have. I think I had to rip the drive out and stick it in another machine to fix it.
~/realm_control_production$ ./realmctl.py stop --no-warn
Oh shit, that is production, not staging. I got the surge of adrenaline and hit Ctrl+C within a couple of seconds.Thankfully it only completed the very first part of the process that prevents new users from logging in and puts up the maintenance page. No playing users were kicked off the game servers. All I had to do was run
./realmctl.py allow
to get new logins to start working again.This is what taught us that we need to have an extra confirmation for actions on the production realm.
~/realm_control_production$ ./realmctl.py stop --no-warn --no-really --i-know-its-production rm -rf .*
Would you expect this to go UPWARD? I would have never. I.e. even if you're in /x/y/z it will traverse up the tree and start deleting /x and eventually /.These days, any time I do an rm with a wildcard, I always prepend with a directory, no matter how trivial, just to keep in the habit. 'rm ../mydir/.*'
Tech: "hey phil, I just broke $bigcustomer's database"
Phil: "..."
Tech: "I typed drop database ImportantDB;2"
Tech: "instead of drop database Importantdb2;"
Even worse because this was lets say.. a very undersized install due to customer cheapness. And all we had for backups (since who wants to PAY for backups?!?!) was a day-old mysqldump off a slave. It took multiple days to re-import all that data. Customer was not pleased. On behalf of all of Joyent, we are extremely sorry for this outage, and the severe inconvenience it may have caused to you, and your customers.
I hate it when people apologize saying maybe it was inconvenient. If someone is using your service and you fucked up, it is an inconvenience.https://gigaom.com/2013/12/02/slap-fight-in-node-js-land/
Either way, it was just shocking to see how quickly they were willing to throw a contributor under the bus who put more than all of Joyent combined into libuv without even trying to understand the his reasoning.