Night of the Runbooks: a DevOps horror story
bitfieldconsulting.com
bitfieldconsulting.com
It's also silly because if the problems are known this well the runbook will only be this good for a moment while the devs are working on fixing the bugs and the runbook will need to be updated constantly after each release.
This requires more testing than most are willing to do and for dev ops to still always be on call to find the new bugs first and document them this way while waiting for devs to fix them and for management to always reserve time to keep this updated.
This means there's a long enough release cycle for this to ever happen which is also unlikely.
The human's role in an ops situation is like that of the "safety driver" in self driving cars: do nothing most of the time, deal with emergencies, and take the blame for the failure of the system.
What you should have is hintbooks. For each alert, have something that identifies why you made the alert. What dashboards to check, what logs to read, what else might be going wrong.
That said, there are still a ton of solutions where "playbook" is kind of the right call. "DB is out of Disk Space" is probably "Increase the size of the disk reservation" for most modern cloud-native companies. They rarely rise to the level of detail that proponents of "Playbooks" suggest, though. How do you do that? Is it clicking through the UI? Do you have the CLI command prepared? Have you thought about why disk space usage has grown suddenly? This is the sort of thing that may be a no brainer you were expecting (in which case - Why didn't you complete the operation in-hours?) or it could be the "That's funny[1]" moment that precedes a major discovery.
Nah. It's not unreasonable. The oddest thing was the feature that would save state and logs in a document somewhere. Logs and dashboards already allow you to focus on certain times, so no need to store that state somewhere else. In fact, that would be worse as it would discourage people from digging further.
> the runbook will need to be updated constantly after each release.
Nah. Nothing in there was unique or special about that specific release outside of the failover. And even the failover is general enough to be a failover. Everything else is standardized to the point of being boring.
> This means there's a long enough release cycle for this to ever happen which is also unlikely.
Daily.
Listen, nothing here was crazy. I've lived this. It's not hard. It's just boring conventions and consistency.
Again, literally the strangest part was saving the state and logs in some extra document.
The other weird thing was not checking the release logs (The release time was weird, too, but I'll ignore that). Rolling that release back would have been my first thought.
And none of this is difficult.
I realize not everyone is going to have the same experience, but it's not a dream, and thinking it's only a dream ignores how within-reach it really is.
Runbooks are supposed to give you a clear path during an incident so you don't run down rabbit holes, but you still shouldn't be running them without understanding what the impact of their steps actually is.
I thought the horror part was that some random person, who doesn't know the system, was doing all this stuff without just contacting the people who do know it.
If the runbook was sufficient to handle it, great! The proper team will have just as easy a time doing the steps, and not be mucking around in something they don't own and won't be the ones to fix when the runbook process itself breaks something because Error A is not Error J, but un-familiarized person can't tell the difference.
Did I read it wrong?
It's one thing to allow anyone to *call* for a rollback of a deployment. It's another thing entirely to allow anyone to, independently and without sign-offs, *perform* a rollback.
I get that the point of the article is to sell the value of run books, which are amazing, but are people really going on-call for the first time without any training or knowledge of the services they'll be supporting? Even with perfect run books that seems unwise.
But I work in e-commerce, not life-critical systems of any sort..
The churn in that team is high that you can’t really give them any significant training on every system. Even a 20 minute video would take weeks to onboard for a system they’ll likely never touch.
As far as I’m concerned a run book is there to stop me from being needlessly escalated to at 3am. If there’s a genuine problem then sure, I or another SME need to be contacted. Most of the time there’s a problem it is a reduction in resilience (so no need for out of hours contact), or something out of our control (loss of power in a location, failure of a circuit, in which case the 24 hour team can raise with the appropriate people.
I can count on one hand the number of problems worthy of a call out, and that’s in the last 7 years.
As such the first entry in my run book is “judging importance”.
Like, if the error is something technical - something is upset and refuses to talk with the database now, that's something you can identify and handle with minimal instructions. I'm working on a side-project at work to automatically generate dashboards for these technical foundations of services in a central discoverable place. Or, is the error something functional? Like, payments from Suaheli aren't being taxed right. There, a generic on-call has no chance. But why the hell would this not be caught in tests, or emerge at night out of office ours, and outside of a planned maintenance if you need that?
And then you're getting to the quality and if your application is a humble workhorse or a feisty goat. With a lot of our systems, we've pushed it to the point so in doubt you can just restart an instance and it most likely will end up better. This makes on-call somewhat chill - in most cases, you just keep kicking the system one or two times overnight until more knowledgeable people come online in the morning and dump the problem in their lap. This includes out databases, just for the record.
On the other hand, I've worked also with enough systems which need like 15 steps to put them back together after a restart."Oh, you need to stop this first, then restart that, then run this jenkins job, then start this first thing, then run this other jenkins job - it'll fail, but that's normal, ..." In such a place, some kind of generic on-call would be impossible. until you get these systems under control.
I still get calls about some service or application I have never heard of.
Technical debt is real and I lose a little bit of my life every on-call week.
And that is, being able to troubleshoot complex systems when you have no idea what is going on, under time pressure.
Every site is a unique combination of side effects that no one on the internet — or even within the company — have any idea on. Yet there is generally a way to start approaching how to figure out what is going on. Though experience with the components of the system helps a lot.
That sort of thing is literally the only thing I miss about the office.
Intuition does play a good part of it for me, though I have worked with good troubleshooters who are more methodical and rely less on intuition for these situations. But then again, having these different takes can help the team as a whole to figure things out and fix things.
Live root-cause analysis of a failure which previously occured in Production - you explain your thought process (check connectivity between System A and B via ping, DNS name resolution, look at open ports, etc...) and they tell u the result like "ping returned 'Name or service not known'".
SWE role in infra.
Yes. Very common
> Even with perfect run books that seems unwise.
Yes
But being oncall doesn't mean that you're supposed to fix every single issue by yourself. You're allowed to ask for help! Monitoring the alerts, taking care of some simple issues (broken CI), trying to fix what you can, and dispatching when you can't. This is usually doable by a new team member.
Even worse, I built and ran the human process that systematically put other devs into that exact same situation. We got good results from the process (as compared to not having the process), but it's obviously not as good as it could possibly be.
No amount of training is enough as features and code gets deployed every day.
It was horrible and I left.
Literally each step was “if error then do automated action”
Alas I have emails which go to a circuit provider when their circuit goes down as they can’t seem to monitor it. It does automated triage, for example checking my router is up and powered, checking the link light on their adva etc.
It tells them there’s no power outage. They inevitably ask the first question. “Is there a power problem”
A run book, or any instructions, is only as good as the person following them.
We is out! No diggity.
in the story our protagonist has no knowledge of the service she is on call for and displays no ability to troubleshoot, reason about, or understand the issue she’s responding to.
instead, she clicks a series of buttons in a runbook to resolve the most basic and most happy-path production issue you’ll ever see.
aaaaand this is about devops how?
I understand the cultural reasons, but I regret the loss of linguistic clarity.
For example, consider the following. "I saw someone get hit by a bus this morning." "Oh, were they okay afterwards?".
"They" is clearly a singular pronoun here.
> I understand the cultural reasons, but I regret the loss of linguistic clarity.
Have you thought about the fact that you're using ignorance of English to complain about a very specific anglophone piece of politics?