I’d say better signal is if sre has its own separate org structure which roots at the cto/vp of eng level. Soon as you have sres reporting to product eng managers it’s beginning of an end.
I’d say better signal is if sre has its own separate org structure which roots at the cto/vp of eng level. Soon as you have sres reporting to product eng managers it’s beginning of an end.
The COO is interested in the stability of ongoing operations whereas the other two roles are about launching new things (Unless I have gotten the definition of COO wrong).
They come to Product from the mountain-top with their holy scriptures (the Google SRE book), prescribing surface solutions to problems with deep roots (pray it away with SLOs).
SRE 2.0 needs to make the embedded model work.
Like, what your asking is "hey, random engineer or engineering manager who is an 'SRE', make the VP of product prioritize your work and staff your team correctly!". That's not something anyone can fix, no matter how good, unless the management is willing to listen, and continue to listen and prioritize appropriately.
This is where a good CTO/CIO with actual leadership qualities will make or break an org, as they should usually be the SRE-esque visionaries, leaders, defenders. When that doesnt happen, and lets face it, it doesnt happen very often, you get cycles of "try this other ops/admin/swe/sre structure and management method" until it falls apart and new management tries something else or they start rolling heads right under the CTO/CIO until it becomes an 1000 and 1 nights. One of the biggest warning signs of this is when tech competent middle managers (worth their weight in gold) start jumping ship enmasse.
> A dedicated model can't work unless the company leadership wants to make it work. Product reliability and performance is not something anyone can fix, no matter how good, unless the management is willing to listen, and continue to listen and prioritize appropriately.
So you have the same answer, incentives and top down buy in. And if you have incentives at management layers, and top down buy in, then youve solved the root people + business problem. And now how/why do you choose between options?
No, the answer is that the SRE org is unilaterally able to stop supporting the product/implement a production freeze if the product org is not doing enough reliability work (or falls out of SLO, or whatever). That's literally the trick, you take the ability to deploy out of the product org's hands and put it into the hands of an organization whose job is reliability, and whose management up to the VP level will back them up on that!
It is far easier to have an organization with the incentive "reliability" and an organization with the incentive "features" and have senior management deal with conflicts when they arise, than to have every middle manager try to correctly split their prioritization in the correct way.
Who exactly does this. And by what mechanism do you enforce:
> take the ability to deploy out of the product org's hands
If it is not "incentives and top down buy in" from above/at both orgs, what is it? Are you literally proposing some reality where VP Alice nonconsensually tells peer VP Bob "you're missing your roadmap ship dates. Im taking away your orgs prod creds because your lines arent up and to the right enough?" Because when there's a conflict between the two it's either going to get resolved by negotiation and common understanding, at and above Alice & Bob, or Bob is going to tell Alice to stay in her lane and get bent.
Edit: For the above, Im interested in the medium/long term. Yes, if you have hived off deployment gates to a different team I believe the could tactically say "no" once. I do not believe they could maintain that position.
Unilateral actions only work when the power dynamics support the person taking action. And those power dynamics are roughly more senior management, broad cultural consensus, and revenue, in my experience. So I'm really curious what enables the dedicated SRE orgs unilateral force of action if its not incentives & top down buy in.
This isn't hypothetical to me. Im a principal systems engineer at AWS who spent years reporting directly to a VP with a billion dollar business, and now work on operations & incident management in the same VPs now larger org running multiple billion dollar businesses. I was there when a billion dollar business took a year off the feature roadmap to focus on availability & performance. No one "told" us to do that, but it was supported as the best thing for the business. I've also stopped working on teams when they weren't making progress with operational investments.
That's one of the key pieces of what makes SRE "SRE" as opposed to an ops team. They are given the power to stop deployment. If they can't, they're not SREs.
> Are you literally proposing some reality where VP Alice nonconsensually tells peer VP Bob "you're missing your roadmap ship dates. Im taking away your orgs prod creds because your lines arent up and to the right enough?"
Yes, and no. No because "roadmap ship dates" aren't reliability. I'm describing a reality where VP Alice nonconsentually tells VP Bob "you've had three outages in the past three weeks, we're out of SLO and so you no longer get to deploy new versions until we have remediations in place to prevent future outages". (and of course they both agreed to this)
Of course, Bob, being a rational human probably starts taking actions after the first outage, and so it never gets to this point, but yes, that Alice can tell Bob "no more deployments" is literally the point.
Google is the only place where this is even discussed. I cannot name any other company that does this. Not Apple, Meta's Production Engineers, Netflix, Amazon's SysDevEngs, Stripe SRE, Uber, Pinterest, and more. This must mean that nobody else does real SRE!
Sweet. Given by whom
> Alice can tell Bob "no more deployments" is literally the point.
Fantastic. Why can she say & enforce this.
You seem to be going a long ways to describe “what” while explicitly ignoring “why” and “how.”
To be clear my thesis is that success in both operating models is built on a combination of top down incentive/buy in and broad consensus expressed as a cultural norm. Im not being intentionally obtuse, but Im yet to see any other hypothesis or method of action whicb would explain the differentiation of ability to execute.
Directly, the answer is the permissions system. In my neck of the woods, devs are quite literally incapable of pushing things to production on their own. The automated tooling handles it 99% of the time, but in the case that they turn it off, only the SREs have production access necessary to actually push new things.
But I think you're speaking organizationally, and the answer there is that that agreement was made by SVPs and is enforced by the PRR process. When you adopt SRE support, you hand over the keys, so to speak.
> Fantastic. Why can she say & enforce this.
Because there's nothing Bob can do except escalate, and SRE leadership will support Alice all the way up to the SVP level.
> To be clear my thesis is that success in both operating models is built on a combination of top down incentive/buy in and broad consensus expressed as a cultural norm.
My point is that the process creates the cultural norm, and enforces everywhere. That is, your org might have these good processes, and that's laudable. But if Andy turned around tomorrow and said "hey, this year your first priority needs to be shipping new features, because we're falling behind", would your org maintain that commitment? Would every exec? The same question goes for every exec speaking to their subordinates.
Its much more difficult to turn to the guy whose job is reliability and reports directly to the CEO and say "hey, your focus this year is shipping features". Like maybe you could say "your focus is shipping features reliably this year", and that should take the focus off of some other things that the SRE org is involved in, but I'm not convinced.
Or briefly, your approach is betting on there being a cohesive company culture of reliability that permeates every team, org, and manager. We both know that's not true, at Amazon or Google or any other company. They're all big enough to have different cultures in different parts. A distinct org and reporting structure, and split responsibilities, ensures that leadership has to be really explicit about such things ("lower your SLOs"), instead of just quietly deprioritizing enforcing or reporting on them. (And I'm generally a fan of multiple individuals with well scoped and distinct responsibility who are charged with reaching consensus over a single individual who makes a decision).
It sounds like you needed a separate chain of command for SREs because your product leadership was incompetent.
If an organization decides that reliability is important (which honestly isn't always the case), then the question becomes whether they can trust Product to include reliability work or not. If not, if reliability is so essential to the business that the business is willing to sacrifice velocity for it, then a separate SRE reporting structure can exist to slow down Product when reliability objectives aren't being met. If reliability is negotiable, then an embedded structure means Product can make its own prioritization decisions and move quickly when it needs to.
The embedded structure doesn't mean that SRE always, necessarily gets overruled. It means that SRE leadership needs observability into when SRE is getting deprioritized by Product, so that consistent failure by Product to respect a core business goal (reliability) can reach executives who will, ahem, clarify for Product that they need to spend some time prioritizing reliability.