Maslow's Hierarchy of Site Reliability Engineering Needs (2015)
plus.google.com
plus.google.com
A business version would be good, something above the pure profit-motive... actualization as a moral, sustainable, environmental, socially engaged organization that gives more than it takes from society. (Note: This is critically different to the SV-dominant self-image band-aid approach of large scale post-profit philanthropy.)
The distinction is practically useful because not everyone is an expert on everything, and if you focus your hiring around very narrow markers- e.g. ability to work with algorithms to solve a coding interview- you may not build an organization that has balanced expertise.
Breaking out the roles and trying to hire for all of them helps ensure you don't get blindspots. The software engineer who spent lots of time thinking about operating systems, Linux internals, and how to build reliable systems out of unreliable parts bring value to the organization in the same way that engineers who focused on mastering algorithms bring value. Everyone needs to be strong in everything, but in practice we have different strengths, and getting people with different strengths together seems to be a good idea.
OS + linux internals + reliable systems = 3 non trivial skills, harder to find and more expensive. The combination is definitely rare.
> Breaking out the roles and trying to hire for all of them helps ensure you don't get blindspots.
If you break out the skills and hire for them independently. You'll end up with a lot of people with the common skills and barely anyone with the rarer skills and the ability to train them to others.
I agree with getting people with different strengths together. It's a NP-complete hiring problem though.
You can't fix unrealistic expectations.
I don't think you want to hire someone for an SRE role who doesn't understand algorthims well enough to do detailed analysis of the codebase, including, say, rudimentary big-O analysis of the code. Good SREs read quite a lot of code and if they're debugging a performance issue and the code contains portions that are obviously exponential or factorial in time, you want them to be able to see and recognize that quickly as a potential cause of the problem being debugged. The same is true of problems caused by bugs or anything else. Likewise, good understanding of what's possible and not possible with software is necessary for area to file (or potentially close) reasonable PRs for the codebase.
The person who thinks in terms of page faults and systems analysis may take a little more time to find the optimal algorithm for an abstract problem than someone who only thinks of abstract problems. But they should know enough about software and computers to recognize the limits of their knowledge, and that's the hard part.
I'm probably being pedantic, but I felt it necessary to highlight the implicit assumption.
(Stolen from Pedro's "Notes from Production Engineering" talk: https://www.youtube.com/watch?v=ugkkza3vKbc , Pyramid shows up around 47:30 min PE = Production Engineer. It's a similar role although not exactly the same as the usual SRE role.)
It's a bit of a different view on the same underlying work. Interesting to see how they slightly differ.
https://docs.google.com/drawings/d/1kshrK2RLkW-XV8enmWZxeRFR...
1. it's a hinderence for people who are studying to become the intended audience.
2. sometimes I'm reading two literatures that use the same acronym, so it's momentarily stupifying.
It really helps when HN posters either clarify things a bit in the title, or if that's not practical, make a comment when you post the link.
[2] I know, ABBR is the new hotness, but I'm still using HTML 4 for my blog and old habits die hard.
Personally I believe this is a direct result of the absolutely massive devaluation of the term "DevOps".
The latter has become a synonym for meaningless mishmash of "everything in operations, and nothing in particular". Using the term SRE is a clear signal that we're talking about engineers with good understanding of systems and operations.
Disclaimer: this distinction is mine, and I have come to it after making the mistake of trying to hire "devops engineers". Put those two words together and your applications pipeline gets crammed with unskilled remote hands who cannot even realise they have delusions of grandeur. It's also - apparently - an invitation for every cold-emailing recruiter to spam you to death.
If you're looking for members of that school of thought, you say SRE. If you're comfortable having a "boiler room" and just need to staff it up, the terms sysadmin and devops work just fine depending on the level of coding involved.
Many more people are paging operators to do DB replica promotions by hand, than are inplementing distributed consensus algorithms and self-healing distributed databases.