When Things Go Wrong
eng.localytics.com
eng.localytics.com
"Why didn't [Engineer X] know enough SQL to understand that query would break the database?"
being bad (While "We need to shut off his/her access." is).
Post mortems have many different levels, technology -> people -> processes -> org -> culture. It's not wrong talking about people, it's wrong to blame them though and stop with people. Asking 'Why?' only stops with company culture.
Usually people make mistakes, they don't want to break things. It doesn't help to not talk about people, but it's important to ask why the person in question acted the way (s)he did. Would others do the same? (most probably) Why? UI issues? Documentation issues? Too many concurrent warnings? etc.
Also important to look from the situation forward (incident happening), not from now backwards (hindsight).
To the SQL question one could ask about the training, reviews, ...
uBlock Origin has prevented the following page from loading:
http://eng.localytics.com/when-things-go-wrong/
Because of the following filter
||localytics.com^
Found in: Malvertising filter list by Disconnect • Basic tracking list by DisconnectThere's really never a situation where I want to download something from a site like that, because I don't trust them at all and I completely disagree with what they are doing.
It's not worth whitelisting - I can read something else.
Have you noticed anything that you'd change? I actually like slack for the most part, but I still find some UX smells in it.
And I actually use it for incident logging too.. never really thought about it but it's one of my favorite uses for it - and it helps the client on that job see the response times and get a general overview.
We have a similar process for triage, recovery, and review, mainly utilizing Slack for our chatops. We use Slack, GitHub and Trello as our three main teamwork tools.
But I have _not_ found chat history to a convenient record of what happened. How do you go back to _just_ show a specific event in history? Also, when there is serious downtime, the channel gets flooded with engineer chatter, and it's difficult to quickly comprehend the root issue, decipher what has been attempted, and know how to help.
In addition to our synchronous chatter in Slack during an event, we open a GitHub issue where we post periodic updates about specific things that have been attempted to fix the problem, graphs from server activity during the event, etc. We'll use checklists to keep track of things that still need to be investigated and what has been done. Sometimes, yes, you need to post in both the GitHub issue and the Slack channel, but it provides a much more digestible, high-level view of an outage.
Additionally, the GitHub issue serves as an ongoing record where you can post the content of the blameless post-mortem, reference follow-on work that was done to harden the infrastructure, etc.
Finally, since the GitHub issue is labelled as "Downtime", it's _actually_ convenient to scan through open and past issues and learn from past mistakes and see what can still be improved.