Google's SRE Book (2017)
sre.google
sre.google
Instead of having feature developers feeling like they have no say in operational requirements, and instead of having reliability staff fighting unstable mess: properly making the contract means everyone gets heard.
Contrasted to devops which despite coming out later was in vogue when this book came out; which caused the muddying of the role of sysadmin to meaning either:
* a sysadmin practicing agile (the original definition btw)
* a software engineer with enough OS skills to carry the pager (the popular one), or
* a team consisting of sysadmins and software developers with no barrier between them (10+ deploys per day style).
Everyone had their own definition of DevOps. So when SRE came clear: sysadmins are needed, stop trying to push everything into one person, heres how we fix the tension between teams: it was a breath of fresh air.
The only revisionist history (that even google seems to forget) is that sysadmins could indeed write code, though it wasnt pretty and didnt have the nice things like mocks and tests. This has changed a little since 2010 at least but it is still dire, even with Cloud making things much easier.
*EDIT:* I've gone from +4 to 0 points in a very, very rapid amount of time. If I have offended you; how?
* SRE are sysadmins with Go
* DevOps are sysadmins with Python
* Sysadmins are sysadmins with Perl
(stolen from twitter, can't find the source)SRE-SEs on the other hand are SWEs, they just have a focus in systems adjacent software whereas a SRE-SWE is someone who can dig into compiler level issues and optimization. Both write application code, do analysis, and write policy. A sysadmin of today would be out of place on a team like that.
Can you please point to the part of my comment that you think disagrees with this?
A good sysadmin was always doing the same thing as a good SRE.
[0] - https://www.usenix.org/system/files/login/articles/login_jun...
ok, but you don't tho (i'm sure there are exceptions). if you do, you will have issues with promo (ask me how i know) so if you don't care about that then sure. also if you're se, you had to interview to transfer out of sre to a pure swe role. i personally didn't have this particular problem but saw a few folks that did.
We just live on a higher level of abstraction and have better tools & processes now.
For example, "developer" is considered offensive, because for some people it's very important to be called "software engineer".
Really good developers don't care about titles.
They don't have time to worry about such, or they have so much money / experience that even if you call them "smart monkeys" they'll be happy with it.
Same goes with sysadmins, SREs, devops, or whatever role you choose.
For some people they have shitty jobs: they don't have such recognition (whether for a good reason or not).
No recognition from work, no recognition from colleagues, no recognition financially, etc, that, if you remove them the title / prestige, obviously they would feel bad.
Source: my experience in a school calling itself "engineering school", and all other schools calling it a "place where to pee code"
That's about the money, not about being good at your work. Ask anyone on the street if you can call them a rat in exchange for a million dollar salary and they'll say yes. It's quite simple.
-shrug-. That's what I feel like in my org at least.
The two jobs are nothing alike, at all, whatsoever.
Sysadmins are support roles. Their functional role is to provide a healthy substrate to run the application layer on top of.
SREs work at the application layer itself. If the system can't scale due to internal architecture, an SRE would be expected to propose a new, scalable design. That would be in addition to maintaining the substrate.
To be clear, there is also nothing inferior about performing a support role. No org can succeed without support.
But the two roles are not the same, and if a job's set of responsibilities don't include shared ownership over application layer architecture, then it can be a great job but it's not an SRE role.
I think one of the key differences to highlight is that usually SREs are engaged early in the design process of new features, and are often driving their own feature changes to the product for reliability or scalability reasons. Those aren't responsibilities or expectations that I've really ever seen in the context of a sysadmin.
I think people like Evi Nemeth, Tom Limoncelli, Æleen Frisch or David Blank-Edelman would have been the equivalent of Distinguished Sysadmins at the time. But they weren't at startups. The places that needed that level were universities, research facilities, telecommunications companies, and the like.
I was fortunate to work under Geoff Halprin early in my career, who while not as well known as those names was a SAGE and USENIX board member and who definitely planted the "don't just do task, engineer yourself out of a job" seed in me.
Sometimes all you can do is shake your head.
There is some generally useful stuff in there, but it probably fits in a few pages vs a full book.
> One continual challenge Google faces is hiring SREs: not only does SRE compete for the same candidates as the product development hiring pipeline, but the fact that we set the hiring bar so high in terms of both coding and system engineering skills means that our hiring pool is necessarily small.
I was thinking, ok so does this mean the book is completely useless for most companies in the world, since they don't have such standards for hiring people or run DevOps this way? How much of the rest of the book is still applicable?
It was written by titans with the SWE ladder at Google, fairly disconnected from the SRE book.
Lessons Learned from Twenty Years of Site Reliability Engineering
Lessons Learned from Twenty Years of Site Reliability Engineering - https://news.ycombinator.com/item?id=38037141 - Oct 2023 (124 comments)
Google Online SRE Books - https://news.ycombinator.com/item?id=31373170 - May 2022 (11 comments)
What Is ‘Site Reliability Engineering’? - https://news.ycombinator.com/item?id=14153545 - April 2017 (86 comments)
Site Reliability Engineering - https://news.ycombinator.com/item?id=13503161 - Jan 2017 (111 comments)
Notes on Google's Site Reliability Engineering Book - https://news.ycombinator.com/item?id=11474002 - April 2016 (93 comments)
Learning the princples and philosophy conveyed in that book helped me tremendously in my career (as a software engineer). Thanks people at Google for writing and open sourcing it.
Seriously. A lot of the book was influenced by Social SRE who had opinions all out of proportion to their own importance and success. At the time, there was some doubt about whether Social's pet theories belonged in the book, considering the varying practices and beliefs of other SRE groups supporting products that people actually use.
This is related to my rule that anybody can title their doc "Best Practices" even if nobody subscribes to them.
Example: the section on backend subsetting in distributed systems is not current. If you wanted the current Google practice you need to read "Reinventing Backend Subsetting at Google"[1], and there are other interesting publications from other organizations.
More practically, I don’t think the book is as useful, as it generally only makes sense when you reach a certain scale that few organizations ever do (imo).
However, we are heading into a future where computing will be everywhere and sensors in everything so in maybe a decade even the “smallest” of organizations may be responsible for large scale distributed systems and operating that would require concepts that are provided in the book.
If you want more about system design and how to design reliability, I suggest reading https://google.github.io/building-secure-and-reliable-system...
learn from it, don't copy from it.
Supposedly, the ratio of SRE to product eng had been growing slowly over the years. The KR to "readjust" that ratio was to bring it back in line with historical norms, i.e., to ensure that SRE continued to scale sub-linearly with SWE/systems. This had (primarily) two facets.
First, it gave SRE teams an effectively-blank check to reevaluate their existing dev engagements and jettison the ones that weren't working well.
Second, it pushed to eliminate old tools/systems/platforms and converge onto the more modern stuff, like Annealing [1]. Fewer crufty platforms means fewer teams needed to run them, and improvements in those platforms have broad impact.
Anecdotally, my own sub-org (within SRE) is growing at the moment. Not by a huge amount, but growing nonetheless.
[1]: https://www.usenix.org/publications/loginonline/prodspec-and...
The rare ones that can do a mediocre job at both (and that won't burn out and switch jobs if told to do both) are usually not capable of doing an excellent job at either.
Using analogies from pretty much any other field shows how dumb it is to combine SRE and SWE, or fuse DevOps (or, god forbid DevSecOps) into one rule:
- Would you have a surgeon drive an ambulance?
- An expert car mechanic manage fleet scheduling and logistics?
- Tell a salesman to design marketing graphics, and have your graphic designer manage high-value customer accounts?
SRE should always be a subtitle for a SWE and not a separate position, and they should always be embedded with SWEs into one team either building products of infrastructure. The shared ownership and toil reduction only works if you have these two things.
All this said, I think the regression is also due to the fact that real SREs are rare. A solid SWE that also has deep systems domain knowledge, understanding how to sift through dashboards and live data, and root cause complex performance problems is a master of many domains and is hard to find.
VERY few companies operate at googles scale. For 99.99% of companies it makes sense to investigate single machine issues.
It's a totally different experience when you have the people who technically own the hardware side of the operations taking no responsibility for the well-being of it, and the people who own the software developing elaborate workarounds for bad machines, and the SREs maintaining blacklists of individual nodes.
What changed?
This sounds... confusing. They moved away from performance?
Both Google and SRE/DevOps have advanced greatly since then, and following the book blindly would be cargo culting.
Edit: apparently this is a controversial opinion?
That's as applicable today as it was then.