Team Structure for Software Reliability Within Organizations
blameless.com
blameless.com
The result? Shifting to this model from an SRE firefighting model has lowered total alerts, incidents, and outages. When we first shifted, people dreaded going on their on call rotation for their team; you expected to be paged multiple times a night. Fast forward and now I can count on one hand the amount of pages I've had on my last multiple on call rotations. Work is prioritized for system reliability because we feel the pain.
During the old days of “you commit code and ops runs it” and during transition to “you build and run it”, we did have some bonuses for being on-call or doing stints in problem management. Now, we bake that expectation into base salary except in locales where we have to split it out to comply with local laws.
Along the way, we've given enough merit increases, market adjustments, and promotions that it's impossible to say whether we had a bump specifically because of the industry's new philosophy on ops. If you wanted to make the case that we did, you could. If you wanted to make the case that we didn't, you equally could.
It also assumes that designing production infrastructure or even systems level thinking is a skill that most software developers have. Certainly, enough do that it's common, but it's not universal.
In this particular case, the org standardized on terraform. Some developers welcomed the ability to self-service. Others wanted nothing to do with it.
Only ever seen it possible to turn around with empowered engineers that have the cycles to fix crippling systemic defects and the business has the means to support it. Otherwise it’s basically a Sisyphean task to maintain things and the business should have offshored the labor by then to save on costs if the cost-feature curve isn’t working out in growth stages.
Microservices merely ensure that the complexity of using different tools per microservice does not lead to an increase of maintenance burden.
Nonetheless, you still have a maintenance burden if every microservice is built upon different tools and processes, since you cannot address patterns of problems easily that occur within combinations of patterns and processes. The cardinality just increases with every tool and process added into the mix.
This burden of maintenance automatically takes its toll on productivity. Teams will not be able to create, maintain or iterate on code that produces business value (as opposed to managing the platform of microservices itself)
I think there is value in a well defined set of processes and tools because there will eventually be platform concerns that become increasingly difficult otherwise.
I do not have insight in how most companies with successful microservice architectures achieved their success, but I would bet my life that a majority of those companies do not let their engineering teams use tools and processes arbitrarily, unless it serves to REALLY produce value that would be unachievable by the tools and processes used up until that point (for example, because performance is usually sufficient, but a specific service suddenly needs to outperform everything that came so far, so you use C, Go or Rust instead of Python)
I am biased into thinking this because I suspect that most of these companies did either start with monolithic architecture (i.e. supporting multiple languages was impossible or just not feasible enough) or they started with a microservice approach focusing on producing value as quickly as possible after an initial ramp up time. Supporting many different tools and processes from the get go would make the ramp up time longer than necessary.
TL;DR Microservice platform maintenance suffers if tools and processes are chosen freely for each microservice, without any push towards unification.
Oof I need to sleep. Don't mind me, just trying to sort myself out.
The crucial feature of microservices is to enable large number of people to work on delivering the same service, but not being constrained by the traditional modes of software release, where the entire monolith could only be updated at a fixed cadence and where the whole thing had to pass through a rigorous suite of tests etc.
One freedom that this gives a team is to choose a different language for their service... on this point there is large consensus. But one need not stop there; you could technically run your service on another cloud with a completely different deployment architecture. All that is asked for this insane freedom is that your service honor the SLA's that you provide. SLA's define the interface, everything else is left to the team.
That being said, shared tooling does provide some nice benefits where learning is shared across teams allowing new ones to bootstrap quickly. However, that is something the team should be allowed to choose.
Commonly we see a team of 20 maintaining 40 microservices, which becomes very difficult to manage and provides little gains. IMHO, this si the wrong approach. Microservices are tiny from the point of view of Netflix, or Amazon.
I've heard managers often tout that "if your code is stable you'll never get paged". But that hides the fact that you're still on the hook to be available. So either you start shirking your responsibility or you reshape your life around some percent of non-working hours being owned by the company still.
Expecting another group to maintain your code seems an awful lot like throwing it over the wall.
SREs have global responsibility to the system and have a correspondingly global view. If your system broke down because of a bug you wrote, that’s on you. If a system three times removed from yours that you didn’t know depended on you broke because you changed a non-API behavior, that’s where SREs shine: they know how to quickly isolate the problem, roll back the necessary systems, and define how to avoid similar problems in the future.
You should have faith in the systems that you build. If you don't want to carry a pager for them, why does someone else want to do it?
During that time I keep my laptop and my phone that can tether with me. I've been on call and at amusement parks. I got paged once and had to go out to the car and work an hour. Calculated risk on my part. What I should have done is ask a team mate if they would cover me for the several hours I was at the park.
However, I do want to work in places where I am responsible for operating what I build. I want to know how well the service I built is doing. If I promise it will be up 99.99% of the time but is not, then I need to know if its down so I can figured out why it went down.
Being on call for systems you build also leads to better software. It makes me design software in ways that are helpful to investigate when it fails. e.g. my error messages provide a lot of context. There is investment in making sure the service is able to terminate gracefully. Metrics are instrumented to locate when things might be off.
What this means is that for things that I build, I often don't get paged as often. Its in my interest to ensure that the service is honest, well designed and well built. When I do get paged, its for something that's seriously wrong.
Anyways, I've been lucky to be in this position. What I've observed for many teams is that they inherit flaky systems which they have to make more reliable; but in the meanwhile every OC shift is an absolute slog, requiring 40 hours/week dedication by people.
If I'm going to be responsible for what I build at all times, then I'm going to be compensated for what it's earning at all times.
That means options, stock, or profit sharing (with a large preference for stock).
Otherwise I feel like this relationship is strictly abusive.
Alternatively, my contracting rate is 150 an hour. I'll agree to work up to 70 hours a week, but everything after 50 is compensated hourly. On call counts, regardless of whether I actually get pinged.
On-call is the type of egregious abuse of employees that an IT union would be able to fight against. Right now companies can take advantage of people who are desperate (H1-B, young parents, etc) and don’t have the freedom to take a stand like the above poster but if all IT workers banded together it would be possible to fight against the practice.
That being said: the SRE model has encouraged orgs to improve observability and set clear expectations and it often coaxes teams into building more reliable systems under the threat of limitless OC toil.
I implemented a similar change at a previous company. Prior to the change, the operations team fronted all on-call issues. After, anything related to the software/services was handled by software teams. We went from nightly alerts to less than monthly. As a bonus, time to diagnose and repair went waayyyyy down due to much better logging.
Note, better logging <> more logging in this case. It meant less spurious errors/warnings, more useful INFO, and better observability of the system.
In my experience (as a chief architect and engineering VP), laying out some baseline metrics closes much of the reliability variance across teams.
As the teams get more competent at baking in the basics (overall load, latency, resource utilization, error rates, error events posted to chat) you can ratchet up the competency to include higher order observables (scaling events, business transactions, circuit statuses, traces, anomalies).
Which is to say its been straightforward (in my experience) for most teams to raise the bar once they know there _is_ a bar and once they can see the bar.
Isn't it best to design for reliability in the first place?
Don't you incentivize that best by having the same people supporting and app as well as developing?
Nothing makes you write reliable software like being the person being called out, or dealing with the customer issue.
Be good with: two or three PM and issue tracker tools, a graphics or UI mockup program or three, various analytics tools and libraries, a payment system or three, several build and packaging systems, a couple reverse proxies, SSL provisioning, several ways to schedule jobs, a few message or event buses, a couple cloud management UIs (command line and browser!) and also WTF all the brand-names in there do, several major database systems with a few different query languages and paradigms, three or more programming languages for weekly use, various communication tools, your own command line environment, the typical environment on your deployment targets and there may be more than one of those, infra as code languages/tools, at least one CI system at any given time, git, the quirks and bullshit of dozens of major libraries across multiple languages, testing frameworks and practices, networking, and on and on.
It’s plainly too much and functionally no-one’s doing all that well, which explains some of what we see from software in the wild. We could probably use about three specialist developers and a secretary for every do-everything developer like that.
However, there is a team that maintains the build system. A team that deals with maintaining the monitoring systems, a team that deals with network, hardware, proxies, etc.
For us, the one we are learning to balance is the data layer. We've traditionally had a DBA team doing most of that. But the team building the service knows their data better. We are shifting to a DBA as a consultant resource model and seeing how that goes.
A mess of a tech stack is more of a problem for developers making future changes than it is for production ( if you have fully automated the deployment, monitoring and service recovery ).
ie the problem of a mess of a tech stack isn't due to being developer led per se, it's a problem of a lack of team working and organisation - which can happen in any team or situation.
[1]: https://www.microsoft.com/en-us/research/publication/the-inf...
In larger, more-established organizations, scrum teams tend to gravitate toward feature development in which it becomes difficult to prioritize reliability work even if the skills are present. With very many feature development teams, it becomes difficult to hire for a programmer that has the primary skill desired and the ancillary skill of reliability and systems operations.
I'm currently of the belief that in most orgs, a team should be dedicated to this. There's the ongoing operations of a service, which I think is a shared duty, and the infrastructure engineering components necessary to abstract the infrastructure sufficiently well that it is uniform across feature-development teams and also reliable, scalable, and secure by default.
As with anything, the devil is in the details.
Some references:
[0] Colm MacCarthaigh, How to take control of systems, big & small, https://www.youtube-nocookie.com/embed/O8xLxNje30M (2018)
[1] Eric Brandwine, Aspirational security, https://www.youtube-nocookie.com/embed/ad9180b4Xew
[2] Marc Brooker, Amazon's approach to building resilient services, https://www.youtube-nocookie.com/embed/KLxwhsJuZ44
[3] Peter Ramensky, Amazon's approach to high-availability deployment, https://www.youtube-nocookie.com/embed/bCgD2bX1LI4
[4] Andy Troutman, Amazon's approach to running service-oriented teams, https://www.youtube-nocookie.com/embed/n1d20Yok000
[5] Colm MacCarthaigh, Amazon's approach to security during development, https://www.youtube-nocookie.com/embed/NeR7FhHqDGQ
[6] Becky Weiss, Amazon's approach to failing successfully, https://www.youtube-nocookie.com/embed/yQiRli2ZPxU
[7] Thomas Blood, Amazon's culture of innovation, https://www.youtube-nocookie.com/embed/2ZQKPUD7vKE
[8] Andy Warfield and Seth Markle, Lessons from Amazon S3's culture of durability, https://www.youtube-nocookie.com/embed/DzRyrvUF-C0
[9] Colm MacCarthaigh, Lessons from Amazon's highest available data-planes, https://www.youtube-nocookie.com/embed/2L1S0zfnIzo
[10] Eric Brandwine, The tension between absolutes & ambiguity in security, https://www.youtube-nocookie.com/embed/GXTvlQXVCOs (2018)
> There are two main types of reliability work. The first is mitigation, which is a linear fix that’s often referred to as firefighting...The second is change management, which is a non-linear fix that proactively reduces the defect rates through projects like migrating to better tools and refactoring spaghetti code. While SREs support both these types of work, they should spend more time on the latter.