Notes on Google's Site Reliability Engineering Book
danluu.com
danluu.com
it's a strange feeling seeing the end of the line for your skills that took a decade of slow grinding to acquire.
If anything, I would argue that the need for the low level languages is disappearing for the general case. As "The Cloud" gets bigger and bigger, and the low level things are handled more and more by a service providers, and they abstract more and more of the day to day, it makes it easier for you to focus on higher level problems.
I do think that the days of being purely ops and needing nothing but shell scripting are going away though. You gotta pick up some python or similar at this point.
this is mostly due to a lack of experience, I guess it's time to really sit down and try to craft some larger projects.
Maybe 5-10 years down the line, but not this year or next.
Without learning some programming you will go the way of the dinosaur, but by learning a bit you can become significantly more effective than you were before.
Additionally, it always takes plenty of time for legacy systems and processes to become forgotten, so even for those who can't or simply don't want to adapt, I feel like there's plenty of work still out there at places where change must occur more slowly, or where long-term investments have been made.
A few thoughts:
> Google places a 50% cap on the amount of “ops” work for SREs: Upper bound. Actual amount of ops work is expected to be much lower
I didn't catch the "upper bound" part from the talk. Good to know! I really enjoy being a developer-who-does-ops. I wouldn't want to be a sysadmin, and 50% ops is probably my limit for happiness.
> I don’t really understand how this is an example of circumventing the dev/ops split
I felt the same way from the Youtube talk. I think there must be a lot behind the SRE role that makes it successful or not: culture, policies, who you hire, how you train, etc. Also I feel like the best sysadmins have been encouraging coding and automation for a long time, e.g. Thomas Limoncelli. But I've certainly been on the "dev" side of the dev-vs-sysadmin fight before, and it makes sense to be seeking ways to improve things.
> Error budget. 100% is the wrong reliability target for basically everything
I think I saw just this month that Google Apps uptime is 99.95%? Some major Google service. I remember in the early 2000s everyone cared about "5 9s", and I feel like for most of us that is just not worth the effort.
> Chubby was so reliable that teams were incorrectly assuming that it would never be down
This reminds me of Nygard's point in Release It! that your theoretical best SLA is the product of your dependencies' SLAs, e.g. 0.999 * 0.999 = 0.998. But in the world of microservices, this logic seems likely to make you underestimate your uptime.
Also I think Feynman's remarks about the Challenger accident apply here: if you are building a new product with, say, 5 microservices, you don't know the reliability of any of them yet. It's dubious to estimate low-frequency events based on "it's hasn't happened yet."
Thanks for sharing your notes. I'm envious you've already got a copy. :-)
https://www.usenix.org/conference/srecon14/technical-session...
As a contrived example, if you've got a microservice that provides data FOO about a request that isn't actually end-user critical, you can mitigate your dependency on it by allowing your top-level request to succeed even if the FOO data is missing. Or maybe you can paper over blips of unavailability with cached data.
But, yes, know what you depend on and how reliable they are, then see if you need to take more action than that if your target is higher than the computed target.
Building reliable services out of unreliable dependencies is a part of what we do. At the lowest level, we're building services out of individual machines that have a relatively high rate of failure, and the same basic principles can be applied at every layer of the stack: make a bunch of copies, and make sure their failure modes are uncorrelated.
I feel like the exact opposite.
What exactly is going to happen when your page doesn't load?
If you're Google or Amazon, the user will try again.
Otherwise the user is going to hit the back button.
My understanding (if I understood correctly), from talking to a friend who is an SRE, is that SREs are also part of the design process. The developers want resources, so they contact and work with SRE teams to make sure their project is both planned for in capacity and can be efficiently served. If it can't be served, maybe another component needs to be deployed that makes the data efficiently usable for the new app or feature (I'm unsure on this, but it sounded like it may have been implicated).
That is, SRE teams become devs of certain components of the project, and work to support the project when in development. This should defeat some of the dev/ops split, because SREs also work on the same project, and are invested in its launch and success.
When it comes to alerting, yes. I've seen it tried many times by competent engineers. The problem is that once you get beyond toy examples into situations with even a mere 10k time series there's so much noise that you can't get any useful signal.
> We could really use something like Outalator, though.
I've not found anything like it yet, unfortunately. Hopefully someone will be inspired to write one by the book, there should be enough detail there to do it.
Magic Systems is anything other than a manually configured (mostly) simple threshold for alerting.
I haven't run into this mindset much at my current job. But in general I think I've been able to lobby for "well, can we at least have a special case that would leave a breadcrumb behind if it does occur?" That way the investigation when it does inevitably occur is swift and there's less debate among ambiguous choices about how to change the design going forward.
I've also found fault injection testing as a great way for disproving statements about what "can never happen."
That said, I've seen the other extreme too -- checking pointers against NULL just prior to dereferencing at every opportunity up and down the stack. In these cases function/module authors succeed only in moving the eventual crash to somewhere far disconnected from the origin of the problem.
Unified, the devops have confidence in what they are deploying and actively work to minimize interruptions via automation.
Source - http://www.retailmenot.com/view/oreilly.com?c=5659596
http://www.amazon.com/dp/149192912X/ref=cm_sw_r_tw_awdm_Tsad...
I tend to avoid ordering books from Amazon too since they charge $5 per book plus $5 per order which usually makes them uncompetitive. Strangely enough I end up ordering most of my books via the Book Depository which is owned by Amazon anyway.
This!
Seems we don't understand the point of SREs at all.
In a world where "ops" and "dev" are split and sysadmins occupy "ops" it is customary that system admins are not programmers, may not know how to program, do not venture into the VCS for the codebase and may not even have rights to check-in code to the software. It would certainly be unusual to see check-ins to the codebase from the ops/sysadmin team.
This leads to the situation where you have 10 year old codebases running on 10 year old frameworks on 10 year old operating systems. The system admins are naturally tearing their f---ing hair out over this situation. The devs, however, have the platform on life support and are off writing new code for shiny systems because that's a lot more interesting and useful than keeping the old garbage on life support. No progress is made and the problem typically doesn't resolve itself until the old systems become sufficiently problematic that the devs rewrite the entire system.
If you have an SRE model it should never get quite this bad. The Devs will support the code in production until they're ready to hand it over for maintenance. When it is handed over the SREs get all the keys to the kingdom and have the rights, responsibility and ability to fix bugs in the software they're running.
If you have legacy codebases run entirely by ops people who don't have any ability to maintain the codebases then you aren't doing SRE.
This is one of the manifestations of the "chinese wall" between ops and dev (which is what "DevOps" and "SREs" are entirely antithetical to--it may be hard to define what those terms /are/ but it is pretty easy to define some patterns that they definitely /are not/). If "Ops" has to come begging to "Dev" to fix their software then you're not doing it right.
Bug fixing is still the responsibility of the developers, which isn't to say that SRE won't help out at times but it's not their role.
> If you have legacy codebases run entirely by ops people who don't have any ability to maintain the codebases then you aren't doing SRE.
An SRE is not a maintenance engineer. Service ownership is always a partnership between SRE and developers. If there's no developers, then there's no SREs.
This is a problem that would be overwhelmingly solved by org-mode in Emacs. Writing the summary in org-mode means you're a 'C-c C-e' away from exporting to HTML or LaTeX.
They hire devs with ops experience.
I've enjoyed what I've read so far in the book and I get it that there is an opportunity here for Google to market the SRE role b/c they are apparently hard to fill but why didn't they put a chapter in the book of what that dev skill set is?
Or else what dev skill are expected of someone working in this capacity.
Seems like lifting the shroud of mystery would do wonders for marketing a role that is hard to fill.
Literally nothing at Google except third-party code (like Linux kernel) is written in straight C. C++ is very common, though.
Python sucks and is useless. Most things that were written in python are now written with Go.
I guess your question is about what happens if a C++ expert ends up on a Java team, but I don't really know. I happen to have landed in C++ world, which suited me.
This is simply not true in any way.
The reason Python tends to suck is not all really nice libraries are ported to it in any reasonable amount of time. (If you need it, you end up doing the work). Go sucks in new and totally different ways.
Edit: Sorry everyone! I totally got my LinkedIn and Google interviews mixed up. What I described above is my LinkedIn experience.
SRE-SE interviews are super heavy on the sysadmin stuff usually, with less (but still significant) attention paid to SWE skills, whereas SRE-SWE interviews may not even have an SRE component (it's possible for candidates in the 'normal' SWE hiring pipeline to be shunted to SRE-SWE post-interview).
While I do tend to spend more of the interview time talking about sysadmin tools, operating systems, networking, databases, security and troubleshooting, I still expect candidates to have reasonably good coding chops.
The difference is that the coding questions tend to be more task-oriented or procedural (i.e. log processing, building automation pipelines, implementing standard unix cli tools, etc.), rather than the algorithmically challenging or math-oriented problems that we'd usually ask SWE candidates.
Both the SE and SWE side SRE candidates need to be able to design and reason about large systems, making trade-offs between performance (especially latency), redundancy and cost.
I guess to a certain extent it must be the luck of the draw - mass log analysis did crop up in my on-site interview as well, so that seems to be something of a theme, although it was more from a high level systems building POV for me.
The Google SRE head honcho says: [1]
>Fundamentally, it’s what happens when you ask a software engineer to design an operations function.
So I think the short answer to your question is yes, they hire software devs and train them [2], but require them to already have some ops/Linux internals knowledge.
[0]: SRE Book
[1]: https://landing.google.com/sre/interview/ben-treynor.html
[2]: https://www.usenix.org/conference/srecon15/program/presentat...
When I was there they hired people with one background and hoped they would be capable of the other. This didn't work out for about half of them. Making things worse, they had no procedure to say, "OK, we hired someone we think is good, but this is clearly the wrong role for them." Which lead to losing a lot of people who had potential.
Making it personal, I interviewed for a pure dev position then was offered a job as an SRE. It turns out that I don't make a good sysadmin. After I left, I was amazed at how many people I met knew of someone else who had had the same experience.
Ironically recruiters see a year as a Google SRE on my resume and think I might want to join SRE teams at other companies. Most of that stopped after I changed my Linked In profile to say that I am never interested in SRE jobs.
I am very good at "tunnel vision". Taking a thread and following it in depth. Completely learning a system in depth so that anything that comes up, I know exactly where to go and what to do.
I am weak on "peripheral vision". Keeping track of 15 balls in the air, and operating with limited information about each. Being effective with limited context about each individual system because there isn't time to learn any of them in depth.
Tunnel vision is a virtue in a software developer. Peripheral vision is essential for a sysadmin.
Now imagine a person with poor peripheral vision trying to learn a system as big and complex as Google's. And being responsible to support 15 different pieces of software written by 15 different teams so that when anything went belly up with any of them, you can trouble-shoot and get it running again.
Nothing particularly bad happened, but I also wasn't accomplishing the job to the standards that Google wanted in that role. And I was not the only person on my team failing in that way.
Ideally it would have been my manager's job to say, "I recognize that this person isn't working out here, is there a better role for them?" That didn't happen. Several months later my manager got fired. I was privately told that my situation was a trigger, but I don't know details.
The fact that my situation is fairly common strongly suggests that the ultimate failure was organizational and systemic. I don't fault Google for hiring developers into their hybrid developer/sysadmin role. I do fault them for not having an explicit onramp/reconsideration process to mitigate the risk that they create by doing so.
In general I believe that Google has a pretty good process. However every process has bugs. And I happen to have encountered one where they pick people who are qualified in one role into a more experimental one, and it only sometimes works out for them.
As others have said, there is a pretty intensive training that you go through and obviously you learn a lot when you are handling things...
[1] http://googleresearch.blogspot.com/2012/07/site-reliability-...
This is a tough position to fill but I've worked with folks who are genuinely good at both, so they're out there.
Sort of like if long-haul truckers needs the skills and training of Fighter pilots.
https://books.google.com/books?id=tYrPCwAAQBAJ&source=gbs_bo...