Simplicity – Google SRE Handbook (2017)
sre.google
sre.google
Boeing is a perfect example of this. I would absolutely read an article proposing principles of engineering reliability from a Boeing eng/QA greybeard. Even as the rest of the company spiraled due to horrible leadership and management practices, many people in engineering and quality control did their damnedest to keep those failures from causing even more harm and loss of life. Those people probably have very valuable lessons to share about how to maintain what quality you can in a deeply hostile environment.
End users making that criticism are confusing the products with the reliability practices.
>regularly does <thing>
I think that might be a reason to suspect that the person doing that is hiding some holes in their argument.
Imagine a smoker telling you to never touch a cigarette
Though there's still hypocrisy in SRE complaining about how SWE builds projects while building equally contrived projects for other SRE to use.
The cause of complexity is not emotional attachment, these are decisions being made. The decision to add feature after feature and punt on maintenance for example is something that has little to do with emotions. There is a lot of agency that engineers, SWE and SRE alike have in shaping how things are. However there can be good reasons to abandon simplicity. The real trouble here is not psychology but that as a profession we are really bad at measuring and estimating the effective cost of maintenance. Part of that is considering measures to improve simplicity and maintainability as cost that comes without gain and somehow less important than features, and then just accept giant rewrite a few years later. A continuous portion of upkeep would likely be more economical and real engineering has always included an aspect of economy - cost vs benefit.
IMHO the loaded accusation of emotional attachment might be rooted in an "us vs them" attitude (SRE vs software engineering) that should have no place in a sober discussion on the value of simplicity and it diminishes an otherwise great text.
> the paragraph starting with "Because engineers are human beings who often form an emotional attachment to their creations, ..." is really out of place.
FWIW I’ve definitely encountered developers clinging to things when the business context has completely changed. I totally recognise the scenario in the original text.
It seems more likely that bounded rationality is at play here, where different parties only know part of the picture (and fail to bring these together and find out what would be best globally.)
Not every case of irrational behavior is caused by emotions though. And when we are making an argument that people are acting against their own interests, it may help to ponder what makes them do so. All the more when we are claiming principles and values that should be accepted by everyone.
"If you don't believe me you are acting irrational / it's because you are emotionally attached" does not seem to be an attitude that gets closer to real causes in a discussion on how to best seek simplicity, but rather a recipe for avoiding discussion or a "thought-terminating cliche."
There must be a better argument for convincing people to let go of code / clean up etc.
Clinging to code you wrote is a very natural thing to do, there are many reasons we do it, and most of the time it’s irrational.
This dynamic comes up often in engineering management.
Separately, cattle vs pets is much older than containers. It got popular with ephemeral EC2 instances when people were first forced to grapple with lifetimes of VMs measured in hours and the ability to scale massively as needed.
When I draw analogies of my past experiences to present situations, that does not mean that my past experiences are the best way to convince people of what is the right thing to do. I still need to do the hard work of pointing out what it is that is in the common interest and why eg deleting stuff and simplifying is good.
In such a discussion it won't help me to say people who disagree with me are generally just emotional, does it? Even if I may have encountered people with such emotional reactions.
I’m not going to just copy and paste the article for you. It’s literally right there.
Because engineers are human beings who often form an emotional attachment to their job security
It's understandably very unwise to admit that Very Complex Solution that cost A Lot Of Money was A Bad Thing
In the same way business is also very reluctant to spend time/money on cleaning up stuff.
I never ever had to make up complex stuff on my own. It always happens on its own.
NIH, CV based development, preference for shiny/new things and a myriad of other "engineer/organizational diseases" exist, you know. And there are even SaaS/PaaS/XaaS marketing teams exploiting such human qualities when making software sales.
I personally think systems evolve the way you describe because of a system of incentives. There are more incentives for features than there exist for refactor and non top priority defect fixes. This comes from the people who hold power to shape incentives and they often do so with conflicting priorities and superficial understandings of the existing incentive structure.
I'd also like to say that it's my own personal theory that systemic issues can only be caused by systemic forces. Individual mindsets cannot be to blame then; if a mindset has become systemic (example: SWEs overly attached to code and features) then your next question should be "why?". There's a system that enforces that, and if you don't look beyond personal obsession then you'll never find it.
When we shift from "reliability" to "safety" we also need to shift from the individual to the system.
This is a great text about considerations everyone operating software services should take to heart.
It applies regardless of if you deploy a monolith or several smaller servers.
If you are only one developer, it might apply in a smaller context.
To be clear, this linked specifically to simplicity which I'm certainly in favor of emphasizing the importance of. But IME the exact opposite happens when people try to imitate Google overall in a smaller setting, where instead too much resources are spent on meta-issues instead of the product being developed.
Nobody is endorsing the practices in TFA “because it’s Google”/in order to be like Google. Sure, people elsewhere make those claims all the time, and they’re wrong, but that’s not in evidence here that I can see.
The article does seem to come pretty close to universally applicable good ideas. Not because of where its author works, but because of the content.
I disagree, I think we can see this time and time a again. YMMV I guess. It's an encouragement to be vigilant for over-engineering when you don't need it because you're not google. I'm not saying that the content is bad, it's a worthy read. Just don't get overeager like the OOP craze phase where would attempt to bend everything into a maze of design pattern because people took whatever books they read way too far. Most of the chapters have YAGNI parts for smaller settings, but it's still worth knowing about what the next steps are.
On the other hand, the book has some nuggets that make it worth reading. But it should be treated as a collection of essays from some very senior SREs rather than a manual.
Definitely! Mostly just a word of caution to not get overeager hoping to apply this everywhere, because "here be dragons".
That's absolutely true, but by design. SRE already had exceptional horizontal knowledge transfer before any book. The book was published (specifically, published) to extend that knowledge transfer outside of Google's own walls so the rest of the industry could also benefit.
I was a SWE-SRE for several years and absorbed a lot just through talks and postmortems. I left Google and joined another big org that was a ... bit ... less far along in its SRE ambitions. I couldn't convince a single person to read the book, even after linking to specific sections to explain why e.g. jitter is important to have in tandem with backoff. Nobody cares, they have boxes to tick and janky bullshit to ship.
Most people are not interested in learning from books these days, inside or outside Google, but at least inside Google you can learn from an unbroken lineage of experienced SREs.
If you're trying to build out an SRE org in your own company, you're better off hiring one ex-Google Senior SRE than you would be if you bought everyone a copy of the book and two weeks of formal training. Actually embedding real-world experience into your teams is the most effective form of knowledge transfer available.
These were also true in the early ages of aviation:
“Perfection is achieved, not when there is nothing more to add, but when there is nothing left to take away.”
― Antoine de Saint-Exupéry
I have not observed this to be the case. After a few revisions there are so much changes that the code cannot be reversed without loosing a lot. A mech aims to cut out the soon-to-be-dead code like a flag is better. But perhaps maybe I’m doing something wrong.
I'm just saying this because Google might be doing this in little islands, not as a company strategy. I don't really know and can only guess from the outside.
What is Google doing in reality?
In most cases to delete code would be a good idea, but to say that source control systems make reverting easier. After a few months most developers will have forgot about those lines and at times uncommenting code & explaining it explicitly might be a better way to preserve knowledge then to rely on digging through GIT.
You build it you run it but may work at their scale
I also find it ironic to see 'Simplicity' touted from the same people who let Kubernetes lose in the wild, but that's a different story for a different time
If you haven't gained any insights from reading that content, maybe it doesn't apply to you or you don't know what you don't know.
> valuable inside the mega corp is actually dangerous outside when people take it as dogma and try to apply it.
mega corp or not, dogmatic principles are usually bad coming from anywhere. The SRE book contains insights that apply to startup, medium-sized companies and mega corps. It's not prescriptive for a reason.
Whether or not Google interprets this advice in a sane way or whether they actually follow it are separate issues, but I think the advice is timely and (at least in my experience) important for many people to hear, regardless of where it’s author works.
The whole point of google inventing a new title and team, from Ben Teynor’s mouth, was that ops should be superseded by a specialization of SWE called SRE.
If your company doesn’t support that, it’s not SRE.
Your argument would be stronger if you could list a few cases like that latest high profile one where GCP deleted some enterprise customer's account. A single one won't cut it for "that's what Google does".
When you’re being paged for the Nth time because of an idiotic problem that you’ve pointed out repeatedly, you too might exhibit these traits.
Pager goes off! Grab pixel. Press finger print reader until it lets me enter my passcode. Ack page. Put down whisky. Shake self. 5 minutes to be logged in and dealing with the problem.
Password. gnubby. password. gnubby. gnubby. gnubby.
Check alert, see playbook, ignore playbook. Check which cell the problem is in. Correlate with rollouts. See a match. Roll back poorly tested dev promo project. Charts recover. Alert not firing.
Log out. Back to whisky.
It's the equivalent of saying "SWE just adds new stubby endpoints to a service, so simple".
The book is exactly the opposite of this. The Principles chapter alone talk about many things that involve actually dealing with numbers (SLO, measuring complexity, etc).
Unless they are testing correlations between these target metrics and business success or some other external cost metric, it’s still “just feels”.
I’ve seen internal crusades against cyclomatic complexity that resulted in massive engineering waste to reduce and reliability saw no improvement.
1. https://arstechnica.com/gadgets/2024/05/google-cloud-acciden...
The problem is that it was never that good. Anyone who has used K8s at scale will tell you at length how it doesn’t scale. People should stop focusing on tech companies like celebrities and focus instead on domain problems related to their business.
Their internal tooling scales just fine, but all it shares with k8s is some of the underlying concepts. Unlike, say, Bazel, gVisor or Gerrit, which are the real thing (minus some secret sauce tied to internal infra). k8s is good software, and best-in-class when it comes to open source options, but the idea that it is "open source Borg" is silly.
No it isn’t, it’s a solution in search of a problem that is needlessly complex, wastes engineering cycles on what could have been product development, and has violated every principal of orthogonal design.
That was a “mistake” that should not have even been possible. If the pension fund had not used a multi cloud strategy the entire business would have been lost. A mistake is not configuring Kafka correctly and losing some data, deleting an entire account should not be given a pass.
“UniSuper, an Australian pension fund that manages $135 billion worth of funds and has 647,000 members, had its entire account wiped out at Google Cloud, including all its backups that were stored on the service. UniSuper thankfully had some backups with a different provider and was able to recover its data, but according to UniSuper's incident log, downtime started May 2, and a full restoration of services didn't happen until May 15.”
Google didn’t recover the data, the customer recovered their data from a different cloud provider.
> This incident did not impact:
> Any other Google Cloud service.
> Any other customer using GCVE or any other Google Cloud service.
> The customer’s other GCVE Private Clouds, Google Account, Orgs, Folders, or Projects.
> The customer’s data backups stored in Google Cloud Storage (GCS) in the same region.
...
> Data backups that were stored in Google Cloud Storage in the same region were not impacted by the deletion, and ... were instrumental in aiding the rapid restoration.
Emphasis mine.
You're quoting, as far as I can tell, an ArsTechnica article that makes unsourced claims about backups being deleted, neither UniSuper's nor Google's previous statements ever mentioned anything about backups being deleted.
I don’t call 13 days a rapid restoration. I also don’t trust Google’s post-mortem documentation more than an independent news organization to be honest about what really happened. Especially while Google is actively gaslighting their users about the errors in its AI search [1].
1. https://www.theverge.com/2024/5/24/24164119/google-ai-overvi...
I'll reiterate, no one involved in the restoration (Unisuper or Google) ever said anything about Google's backups being deleted, in fact basically everything Google and Unisuper have said specific that it was only the VM config that was removed. Ars made up the thing about backups being deleted, which makes an exciting headline, but it doesn't appear at all reliable or based in reporting, just conjecture.
So you just label a reputable news outlet as fake news and then move on..?
“UniSuper had backups in place with an additional service provider. These backups have minimised data loss, and significantly improved the ability of UniSuper and Google Cloud to complete the restoration.”
https://www.unisuper.com.au/about-us/media-centre/2024/a-joi...
Those quotes were pulled directly from UniSuper’s website. Google deleted an account, lost the data, and then took 13 days to recover the pension fund from data stored on another data provider. Maybe you should consider that your employment at Google is damaging your objectivity.
Like I keep saying, nothing in the primary sources supports the claims either that backups or that the accounts were deleted. You've jumped to a particular conclusion, and seem unwilling adjust that conclusion in light of new evidence.
You're making a strong claim, I'm asking you to source it specifically. Instead you're taking a statement from which you can draw multiple conclusions, and picking one (that has been contradicted repeatedly) and telling me I'm unwilling to accept the facts. But they aren't facts, they're your interpretation of vague statements.
I'm happy to accept facts. Facts like "Data backups that were stored in Google Cloud Storage in the same region were not impacted by the deletion" are very easy to understand and difficult to misinterpret. Do you disagree?
Like, even the additional reporting the ars article links to (https://danielcompton.net/google-cloud-unisuper, https://x.com/milesward/status/1792909048830214607?t=Vu__q1h...) basically contradicts both their and your conclusions. You're weirdly hung up on this.
Again, the same quote I already quoted above, direct from UniSuper’s website. They needed to use their backups at a different cloud provider, as GCP’s data wasn’t recoverable. I don’t know why you’re arguing so strongly against this.
Put formally, we have statements that
- 1. A and B exist
- 2. A was used
- 3. A and B were used
Your conclusion from these statements is that, because (2) A was used, therefore B does not exist. Hopefully putting it like this makes it clear why I'm so confused.Maybe if Google focused on doing actual work instead of writing feel good engineering pieces, they wouldn’t have the Google graveyard and an unstable cloud offering that may spontaneously delete multi-billion dollar accounts.
That same sort of thinking is what led to the downfall of yahoo.
Are they?
Are you saying Google’s c suite isn’t out of touch?