Why Blameless Post-Mortems
medium.com
medium.com
Instead, you need to look at the reasons underlying why Dave pressed the button. Is the button confusingly labeled? Is there no playbook on what the process for pressing the button should be? If you fix those underlying problems, you will prevent other people from making the same mistake Dave made.
It sucks being the one who caused an outage, but having a culture where people put blame on the process rather than the person takes away a lot of the stress and naturally motivates people to solve the core issue as quickly as possible. The points in the post about psychological safety and underlying systemic causes of an outage are spot on.
My employer does a reasonable job with this. Our tools usually make the path of least resistance the right thing. We have guardrails that help prevent accidents. But at the end of the day, our controls presume competence and good faith, and there's always an "I know what I'm doing" button. Actions taken after pressing it incur a bit more personal responsibility.
When your employer does the same thing, it leads to “good tools with sensible affordances and guard rails” for competent responsible people.
Why is that?
A finding of "the process was reasonable and the individual wasn't" or "a process guaranteed to stop this kind of action would be unreasonably heavy" may sometimes need to be on the table.
Also, a good postmortem process ought to treat "Why didn't we do this in time" as a problem to solve, too. (To be fair I don't immediately know of good examples of this.)
> explaining to him why he shouldn't press the button won't prevent anyone else from making the same mistake.
How I interpret the point is: don't fix people, fix tool.
You can't tell someone not to break the system, you need to make the system less easy to break, labels on buttons, confirmation dialogs, no double negatives, approvals if need be.
Reminds me of a deployment with a customer (some 15y ago), I wasn't allowed to "touch prod" myself so I was sitting next to a button pusher doing what I was telling him to do (irony...). At some point there was some trickery involved to reboot HA instances one at a time rather than simultaneously, and you had to press "No" when prompted if you wanted to reboot. Before the dialog showed up, I pre-conditioned the guy, telling him "it's going to ask if you want to reboot, say no. Whatever you do, do not click on yes". This wasn't the first time he deployed. He said "oh, ok" then sure enough clicked on yes, and the IT police was in our room 2 minutes later because of a prod outage we just caused.
I have thought many times of what could have prevented that - allowing me to deploy myself, doing more training to that guy, finding a competent replacement to him, more documentation. There is no process that would have worked better to fix the issue than finding a way to remove this dialog box - today that would be done through automation.
I like to think SWEs are in creative businesses... But this doesn't always hold true...
The whole point is CYA doesn't work. Its not trying to make it idiot proof, it's trying to remove the idiot from being involved at all.
Despite being a strong believer in it, I've worked at places that did blameless post-mortems and it was sometimes frustrating because remediation often felt like a second-class citizen. Remediation is often much more difficult in these environments because systematic changes can obviously much more complex than "fixing" an individual's behavior.
If someone pushes the big red button accidentally I'm fully on-board with the "well what led them to push it and why were there no safeguards?" but it also kinda sucks if you spend so much time on that conversation that there's no time left for "ok, who's going install the confirmation screen for the big red button?"
Sometimes the "remediation" was even just "well, that's just the way it is and we can't do anything about it easily, but hopefully the post-mortem itself will serve as a way to inform people not to push the big red button" which is obviously frustrating because memories are short and people come and go quite frequently in the tech industry.
Anyway, and this might be obvious to some, but remediation is important and sometimes incredibly difficult in these environments. Adopting such an environment requires a commitment to actually fixing the systemic issues. It can sometimes feel more frustrating than "blameworthy" post-mortems especially if it's the same individual(s) over and over again.
Maybe I'm missing something, but that seems orthogonal to the blamelessness. Blame won't necessarily bring remediation front-and-center either (and will probably do far more harm do that end).
My favorite "hate" on the blameless post-mortem's is when it comes down to altering a process that someone owns, because then you're not only fighting for time-budget to fix and issue but you're potentially fighting what is, to them, part of their job and existence.
To get tangential, this is also why I hate big-A Agile, because if you have a person who's job is "do Agile," the scrum (master|lord|bag|dumpster), any changes, deviations, or complaints end up feeling like you're attacking that person's job. Never a fight you want to be in.
A person who plays a part in a costly error is going to be much more cautious in the future and be very helpful in improving systems so it can't happen again. You want to foster that approach rather than discourage it.
Unless they're just a fucking cowboy all the time and you should just fire him now because he's going to do it again. Maybe gauge the level of surprise of his collegues. Are they shocked at how this could happen or are they going "sounds like Jerry".
Some people work more carefully than others. This is absolutely a thing in software development/operations (my fields). When I was young, I had more screw-ups because I was still learning "how to work carefully" with certain things... It's hard for me to admit this but it is 100% the truth. When things went wrong, I felt the need to take it on my shoulders vs. to put it on the rest of the team who likely had nothing to do with the error.
As with all things this isn't black and white, which is why I'm making this comment.
I also favour responding to mistakes quickly over preventing them completely.
I've seen a lot of teams where there was the one keeper of the systems who knew how to do all the deployments and migrations. In every situation they became a bottleneck for the team, and prevented any real improvements being made.
Even at my last role people would get angry at me when I'd ask for "command confirmation" and/or "an extra sets of eyes".
In my world, I favor getting it right vs. responding to mistakes. It's the difference between "scalable" and "firefighting".
Having a SPF asset who holds the keys to the entire kingdom is a completely different problem, and a different conversation.
Ha, yea. I can get behind this. There are times that one dumb thing causes the entirety of my tests to crumble and I'm always admittedly laughing saying "that's why we have tests!"
Sometimes, it really is one asshole and they need to be fired. Most of the time, it's not.
But leaders are.
Most enterprise post-mortems I've read will deliver pages and pages explaining (or justifying) poor system architectures. But they will never, ever criticize the executive or tech lead who decided on that structure to begin with.
Everyone makes mistakes, for sure. One stray outage is likely technical in nature. A simple oversight that can be corrected. But if an org's systems are constantly failing or breaking, then it falls to the leader of that org to reshape its structure.
Or they can pretend that blame doesn't exist. The system is fine. A band-aid will suffice.
Without commitment to fix systemic failures, "blamelessness" enables the abdication of leadership.
Certainly explains their popularity lol