Incident management at Google – adventures in SRE-land
cloudplatform.googleblog.com
cloudplatform.googleblog.com
"Can I handle this? What if I can’t?" But then I started to work the problem in front of me, like I was trained to, and I remembered that I don’t need to know everything — there are other people I can call on, and they will answer. I may be on point, but I’m not alone
It might be because I'm currently training people in this realm, and this is one of their biggest fears, or maybe because it was my biggest fear, but its so true. We're a team. We're here to help. At least if your SRE org is any good. Never be afraid to ask for help, and never be afraid to admit you don't know something or it might be outside of your comfort zone.
I'll take willing to learn and readily able to admit knowledge deficits over someone who doesn't any day of the week. Great book they're working on, great article on this. So many gems, but this one stuck out for me, and its pretty relevant to me right now.
1. It encourages everyone to be aware that we're here to help
2. It encourages learning, the number of request for help decrease over time once the experience and familiarity ramp up
3. It exposes everyone to different types of work.
There are exceptions of course. Our P0 bucket has a dedicated set of people that handle that, and they are hand picked because those are house on fire situations that need the experience. Its also the one we put the juniors on when they know the ropes and are ready to take on the critical tasks so they can advance themselves (its good experience for career and personal growth I feel).
The other thing i like about this is that as a manager, I can actively encourage behavior by claiming a ticket, or helping someone with a ticket. I really want that culture effect to happen from the top down.
"Rock stars" are downright dangerous, as are people who prefer to make things up rather than admit ignorance.
A new SRE doesn't need to know everything (and can't). But he absolutely needs to be curious and willing to ask for help.
We look at essentially, (but not unequivocally) these things
1. Proven experience and desire to learn is a must. One of the best i have ever worked with came from a place where they worked pilot projects, and had to manage all their infrastructure themselves as the developers. No certs, no formal SRE experience (IE, thats not why he was employed previously). One of the best. He loved the work, and I could tell he learned so much doing this. It doesn't have to be this extreme, but having a proven interest in this line of work is top priority.
2. Is their depth better than their breadth? It is correct that you need a LOT of breadth, however I value the depth first. I'd rather someone have say, a medium about of breadth on the different technologies out there and a lot of depth on core subjects, like container management (this happens...everywhere nowadays) or cluster management. I don't need you know every single implementation of this in depth though. I need you to at least know one implementation of this in depth. I can build on that.
3. Because of the first 2, I need someone who is team oriented, as always.
> core subjects, like container management (this happens...everywhere nowadays) or cluster management
Curiously, these are subjects which most Google SREs won't know much about. One team deals with all that stuff as a service so the rest of us can get on with something else.
What would I pick out as our core skill sets? Ignoring technology-specific details that won't apply anywhere else: troubleshooting a system that you don't understand (reverse-engineering it as you go), and non-abstract large system design.
Broad experience, depth on a few topics. It's not impossible at all, it just takes time.
(edit: Note this may have changed since I left the company in 2009. It's been a while!)
There is a reason that this is an impossible to fill role.
I myself have deleted the data file of a production mysql server, because I have no idea what I am doing. I need to call teammates in 12am to learn how to take care of the mess I created.
However, having it be policy to ask first then trade or have someone else own it means essentially as a team, we can own that issue with you as you learn. It helps a lot come crunch time.
Mistakes will always happen though. If they are genuine and not from a lack of caring or trying with an earnest/logical thought process behind it, it doesn't matter in the grand scheme. I'd rather have my entire production server go down on an error like that than deal with a situation where the person was just being negligent to their duties and wasn't willing to be humble.
In short, you'll be fine :)
I wrote down some thoughts about this recently: http://zalberico.com/essay/2017/02/21/asking-questions.html
For some things, you should be able to trust others.
It's funny, because it boils down to "don't lie" (to yourself nor to others).
that sounds really nice
As a manager, this made the concept of 20% time a lot more clear. These are people with the knowledge and incentive to build a hierarchy of systems that progressively remove risk from their work. This is in fact their primary business objective. And we need to make sure they have time to do that, vs working them to death with manual remediation. It's a great lesson.
Incidentally, Stackdriver contains a simple alerting and incident management tool that's really nice to use. Hopefully it gets more robust as time goes on and larger and more complex orgs move to their cloud. Edit: not Outalator.
That's not what our 20% time is for, and 20% is way too small a number for that purpose. "20% time" (the way we use the term) is for personal/career growth/scratching itches.
Time spent on building systems that make our service better is my primary job. Manual remediation ("toil") is something to be tracked as a dangerous antipattern that must not be allowed to take over.
Toil and oncall response should be less than 20% of my time, together. At least half my time should go into engineering projects. If the level of toil is in excess of 50% of team activity then I would expect only percussive intervention to get the team out of this situation.
The incident management tool is not Outalator. Outalator is the pager queue management tool. The incident management tool is for manually creating incidents that have much broader visibility than Outlator does.
As someone who has been incident commander a few times, incidents tend to have broader impact beyond your immediate team or owned jobs.
progression of system operations from manual, to scriptable, to automated, to a fourth category I hadn't even known existed: autonomous.
Completely agree. The first "eureka" moment is when you define all your infrastructure in code in something like Terraform. Magically networking, firewall rules, disks, instances, are all provisioned and dependencies calculated. It is quite a breakthrough from running CLI commands or using the web interface to allocate infrastructure.Plug: I wrote a blog post on getting started with Terraform and Google Compute Engine for those interetested https://blog.elasticbyte.net/getting-started-with-terraform-...
What's the difference between automated and autonomous?
I'd highly recommend it if you're in the Ops feild. Probably the best book out there on current large scale Ops practices.
Interesting write-up though
Chaos Monkey is actually taking down production systems to make sure the system as a whole stays up when those individual pieces fall. Google does have (manual, not automatic) exercises doing similar things called DiRT (Disaster Recovery Testing), but it's not related to the SRE training exercise.
(standard disclaimer: Google employee, not speaking for company, all opinions are my own, etc.)
WoMs are a training exercise, intended to build familiarity with systems and how to respond when oncall. A typical WoM format is a few SREs sat in a room, with a designated victim who is pretending to be oncall. The person running the WoM will open with an exchange a bit like this (massively simpified):
"You receive a page with this alert in it showing suddenly elevated rpc errors (link/paste)" "I'm going to look at this console to see if there was just a rollout" "Okay, you see a rollout happened about two minutes before the spike in rpc errors" "I'll roll that back in one location" "rpc errors go back to normal in that location" ...etc
(Depending on the team and quality of simulation available, some of this may be replaced with actual historical monitoring data or simulated broken systems)
The "chaos monkey" tool, as I understand it, is intended to maintain a minimum level of failure in order to make sure that failure cases are exercised. I've never been on a team which needed one of those: at sufficient scale and development velocity, the baseline rate of naturally occurring failures is already high enough. We do have some tools like that, but they're more commonly used by the dev teams during testing (where the test environment won't be big enough to naturally experience all the failures that happen in production).
Disclaimer: I work at Google and ran the DiRT team for a few years incl. incident management itself.