De-siloing incident management: make reliability engineering everyone's job
rootly.io
rootly.io
> SREs aren’t necessarily organized as part of development or IT ops teams.
SRE teams should be comprised of software engineers. Who do software engineering - specifically to tackle the thus-far intractable problem of managing operational complexity for the SDLC. It says as much in the first chapters of the SRE book. The idea that SREs are not engineers, or developers, is nuts to me. Of course, most SREs _arent_ developers, they haven't ever been in that role. Thats fine, but organizations need to understand that transposing Sysadmin into SRE doesn't magically change anything.
> No role in CI/CD for reliability engineering
Phew boy - this really underscores the weakness of the SRE meme, in my opinion; and highlights the usefulness of the Platform Engineering meme. CI/CD is the fulcrum where effective SREs/Platform Engineers can insert a _ton_ of value into the business, and if they're not allowed to have a leading role in its care and feeding then they are basically being kept from doing their jobs. I've seen this many times over. Another pathology is an organization where CI/CD is neglected in favor of shinier, more whiz-bang toys - this is just as unfortunate.
Just my two cents - but as I always say, everthing in moderation except moderation.
t. SRE/Platform Engineer/DevOps for 6ish years
https://charity.wtf/2018/10/24/ten-platform-commandments/
https://martinfowler.com/articles/talk-about-platforms.html
https://lethain.com/pierceable-abstractions/
https://srvaroa.github.io/paas/infrastructure/platform/kuber...
Hope that helps - had to dash this off fairly quickly.
For those who currently care about X and they don't get enough support in their organisation, it makes perfect sense. For the rest, nothing needs to change. After all those other guys should have it.
The article issues hypothetical directives "Include all teams in testing", without any consideration as to what kind of activity might make this feasible. All teams simply reply "no thank you". And it's back to square one.
To be clear, I am not saying these suggestions are wrong. Just naive.
On the other side of the spectrum, the main problem with "Everyone owns X" type initiatives is that often nobody gets measured on X but they do get measured on their other responsibilities. The predictable result is that all those smart and driven people you hired will realize that spending time on X will not get them promoted but neglecting it does not cost anything.
But this can get a little over-complicated, and still not quite specific enough. You adopt these 4 practical strategies, and what do you get? Blind adherence to a rule. Empty formalism. Following the letter and not the spirit. And the reason is, all of this shit is hard! It's complicated and requires deep knowledge. Asking a random person to do a thing they don't really understand well isn't going to result in overall better outcomes.
I think the answer is perhaps the hardest one: we need better education. We need an actual school where people go to learn about every facet of everyone else's job in a tech shop. If everyone has the deep knowledge of how each other's jobs work, what their considerations are, and what needs to get done when, we don't have to struggle putting processes in place to constrain the eventual chaos of competing priorities and unknown procedures.
Put simply: we all need to skill up.
What does their problem SRE team look like, and what does their ideal SRE team look like. In my experience the SRE team is doing things like building the CI/CD platform, the Environments, Infrastructure, etc. So that when I make an application, and commit my code, it then creates the infrastructure, and deploys my code. In this paradigm, an incident with my application goes to me, and an incident with the infrastructure breaking is on the SRE team. An incident with a firewall being configured wrong doesn't make sense to go to the UI devs that put a new microservice up, and an incident with UI not displaying values right doesn't go to SRE.
Since this article says there is no CI/CD in the SRE team, it makes me think that this is some kind of operations incident response team, and thus I would say misnamed.
A dev moving to sre will often need to be taught the importance of documentation, how to have smaller immediate impacts and a customer service mindset.
Theres a middle ground that is difficult to find, but is essential to balance the value you bring to the company with the amount of future work you're making for yourself.
Some of the worst sre I've seen involve a dev taking the title and refusing to support, collaborate, or communicate. Alongside a former sysadmin that refused to document, bypassed ci/cd process, and constantly declared tasks as impossible.
When I'm interviewing a sre candidate I focus on identifying if this person has behaviors that make them Ill suited for sre. Technical skill is often easily faked, and your behaviors around how you admit you don't know are more telling than you not knowing.