Handling Human Error in the Datacenter
spiteful.com
spiteful.com
on all servers that I have physical access to, I use zipties on both ends of the power cable. you need a knife to unplug anything. One problem, at least, solved.
Generally, I categorize mistakes as 'mistakes of knowledge' (that is, I did the wrong thing because I believed something that was incorrect.) and 'mistakes of inattention' (where I knew it was the wrong thing to do, but I wasn't paying attention and did it anyhow.)
Generally, you don't make the same mistake of knowledge twice, so I don't worry about them much. They happen, but they only happen once. Learning, we call it.
Mistakes of inattention are much worse, in my opinion. without further action, I will almost certainly repeat a mistake of inattention.
The idea is that every time you make a 'mistake of inattention' you put in place a procedure that will prevent the mistake.
What you describe is the operations equivalent of a software guy hacking a one line patch around a bug. Too many of those without any cleanup/refactoring and you start introducing new bugs while fixing old ones.
the alternate is to just fire people when they make dumb mistakes... but for me, that is unjustifiably expensive, especially when compared with the cost of adding process.
Sure, if you control the app you can set things up such that losing one physical server isn't a problem, but that is usually a lot more expensive than a few zipties. One thing I've learned in my years as a computer janitor is that a fix I can put in place now with materials I have on hand gets done. A fix I suggest to the devs, even if they agree it is a good idea, rarely gets done.
Process, when it includes humans, is bureaucracy. And yeah, you need to make sure it doesn't get ridiculous, however, having a checklist for, say, installing a server, or replacing a bad disk is still, I believe, in the realm of 'good bureaucracy.' When you get paged at 4am after a late party, you can't count on being sharp and remembering everything. You have to accept that especially for emergency maintenance, sometimes your admins are not at 100%.you also need to accept that things you don't expect will happen. You can't always rely on process. I'm just saying that you can use process to avoid making the same mistake twice.
I don't want people trying to pop zip ties with knives near my cables carrying production traffic.
I had a script on my server that did clustering of stories from different news sources. The script also had some test methods which deleted all the clustered data and rebuilt it. I once accidentally ran the "cleanup method" on prod server and that created disaster because somehow cascaded deletion took place. I had to refer to replay log to get everything back and took hours of efforts plus a lot of pressure. From then onwards I placed a check on each of my script to get a confirmation twice before executing any such test method on prod server.