Founding Uber SRE
lethain.com
lethain.com
> The name was unchanged, but it was soon a very different team. That unchanged name concealed the beginning of a significant leadership and cultural shift
Unfortunately a very common problem. New person weighs heavily towards their own great experience previously, fails to understand new org and its problems.
One of the best infrastructure managers I worked with shared their approach with me: the first 6 months, they would go to all the team leaders for the teams they served and listen to the problems they faced. No solutions were proposed. Just understand the pain points first, develop solutions after learning those.
I often wonder if we should have social scientists study these phenomena and create some guidance so people don’t keep repeating these mistakes.
Businesspeople have more to learn from their own than from people in a completely different sector, just as artists like painters and sculptors learn more from their own than art historians and theorists. Much of academia is masturbatory, of no use or interest to people outside the specific sub field of academia it comes from, and the further divorced a field is from use by or contact with practitioners or consumer demand the greater the proportion.
More bluntly most business professors never even try to become successful entrepreneurs. They’re unwilling to bear the costs of the market test because they don’t expect the benefits to be worth it.
> Deming edited a series of lectures delivered by Shewhart at USDA, Statistical Method from the Viewpoint of Quality Control, into a book published in 1939. One reason he learned so much from Shewhart, Deming remarked in a videotaped interview, was that, while brilliant, Shewhart had an "uncanny ability to make things difficult." Deming thus spent a great deal of time both copying Shewhart's ideas and devising ways to present them with his own twist.[16]
> Deming developed the sampling techniques that were used for the first time during the 1940 U.S. Census, formulating the Deming-Stephan algorithm for iterative proportional fitting in the process.[17] During World War II, Deming was a member of the five-man Emergency Technical Committee. He worked with H.F. Dodge, A.G. Ashcroft, Leslie E. Simon, R.E. Wareham, and John Gaillard in the compilation of the American War Standards (American Standards Association Z1.1–3 published in 1942)[18] and taught SPC techniques to workers engaged in wartime production. Statistical methods were widely applied during World War II, but faded into disuse a few years later in the face of huge overseas demand for American mass-produced products.
Of course it’s difficult to know how much management change was due to him. E.g. I’ve seen claims that he influenced the re-building of Japanese industry after World War 2, but I don’t know if it’s true.
The British Library has interesting pages about Drucker and other people they describe as “management thinkers”: https://www.bl.uk/business-and-management/management-thinker...
I once got rejected from a manager interview with a billion dollar unicorn. I said I would spend my first 3 months understanding and observing. The director retorted and asked rhetorically what if my team complained that I wasn’t doing any work?
It left a bad taste in my mouth. And I have pondered for months whether I mishandled the question, fundamentally misunderstood management, or whether this unicorn is doomed to fail-at-scale because of poor organizational intelligence.
[1] There a few times where big, shocking changes are necessary to reset a team in a bad situation. Most of the time this is not the case.
There’s a lot of people out there who are smart and can perform well in an interview, but aren’t do-ers. They don’t make things happen. You don’t want to give off any vibes that you are that kind of person if you are interviewing for a fast-moving company.
I believe Amazon calls it “bias for action”.
You should certainly look to understand large problems, but if you know a quick solution to a simple problem - just implement it.
Whether or not customers asked for it, or whether the same app design fundamentally could run on a single machine under their desk, doesn't really come into it. And bootstrapped companies (Instagram or Stackoverflow for example) often do manage to run on the single machine under a desk.
As a leader your job is to enable those people. Then as time goes by you'll get your own knowledge of the problem space and can start taking some initiatives yourself too. But there's no need to literally do nothing for so long.
People often misunderstand "empower people " with "I'm a dunce and can't help anything for months on end" as if magically on month 3 you're now enlightened. Perhaps the interviewer was checking if you understood the nuances of these situations.
Thanks for sharing this perspective.
Instead you could have said how you'd work to create a culture based on your values of x, y and z. Ideally your values would align with the corporate culture and would motivate and inspire team members. You could have even told the director you'd have held meetings telling your team that you wanted to find out what their problems were. Tell them that on your watch things will get better, etc. Reassure them and provide a vision.
You don't only need to observe to discover problems, if you only ask people are likely to tell you, especially if there's new blood that might be able to create changes.
This is the heat death of any organization. Organizations can absorb a stunning amount of dysfunction and somehow keep ticking.
> The director retorted and asked rhetorically what if my team complained that I wasn’t doing any work?
I think you can interpret that question two different ways, and it's hard to tell which one it was without being in the room/being in their head. The more generous way to interpret this is that they were checking your understanding of the role a manager plays outside of pushing changes and decisions.
That said, if it was clearly the director communicating that a manager should be jumping in and imposing a vision immediately, then you would've been misaligned philosophically, and that's probably not a job you'd want even if you're qualified. Being a manager while fundamentally disagreeing with your own boss is having one hand tied behind your back.
It's possible that the interviewer was of the "Jack Welch mindset" and worships action and making snap-decisions-- Captain Kirk style. A LOT of people see that as the ultimate model behavior for an executive. Or maybe, he instead asked that to see how you would push back to the inevitable pressure from others in the org that wouldn't respect someone that's not biased for action?
Let them complain so you know who to get rid of in 3 months.
That three months buys you credibility to impose bigger changes.
Just going to the place where the thing happens and looking at the thing happening to better understand is a hugely undervalued job.
I don't want a detailed digital report of how my child is doing in preschool. If I want more details, I'll take a couple hours of vacation and join them in the preschool, just to observe. I don't need a product manager to tell me how a customer uses a product in detail. If I want to know in detail, I'll visit them and watch.
Instead in the corporate world we have layers of layers of management demanding progress reports and skill assessments and performance rankings and velocity audits and what have you. I'm sure this eats way more time than if they just spent a couple of hours every now and then working with and asking things of the people who do the actual work.
This rings very true in my experience. Every new layer, every new TPM organization adds more overhead. They start off with “were gonna boost productivity and get things done faster!” but when that doesn’t happen the result is a demand of more process, more bureaucracy, more JIRA/story point blah blah all of which ultimately falls on the engineers who still have to do the actual work but now have to learn how to bullshit the bureaucracy.
That might sound obvious, but it's not always. Having commonly used tools on a central toolboard instead of at each workbench makes it seem like part of the task to go to the toolboard, when in practise it's an irrelevant wasteful side-task. You can improve how smoothly (quickly) you go to the toolboard, but the improvement in speed comes from moving the tools to the workbenches again.
Likewise, smoothly doing paperwork of the sort we're discussing here will make your paperwork faster but it won't improve your first-order work.
> It's something one ought to do all the time, to better serve the people one is there to serve.
IME, it seems like most managers believe that you are supposed to serve them, which I think gets right to the heart of whats wrong in many organizations.
I started working at Uber the same month as Will after the infrastructure team passed on my resume without even a phone call and handed it off to another infra team which Will neglected to mention -- a "shadow infrastructure" team created as an artifact of technical differences and personal grudges which both doubled every six months like our head count, we were responsible for edge services' business logic written in Node.js instead of the Python 2 predominant in the rest of the engineering org. I joined as a junior engineer out of two failed ruby on rails startups with ~0 positive Node.js experience but a good enough sense of Linux systems engineering and administration to be doing release engineering and on call for those edge services three weeks later. Everyone had a similar trial by fire in that time, and many didn't "make it" a month despite having a smoother hiring experience than the story Will related.
The infrastructure org was lucky to have Will. I didn't have a manager for months, reporting to a 20-something "director" who was too far in over his head and too up his own ass to care about that, and taking day to day "marching orders" from someone who is the most technically brilliant people I ever worked with but with the EQ of a bat. This leadership debt was a recurring pattern through out the whole engineering organization throughout the years of insane growth, even after the "new SRE" came to be. But that toxic division and leadership debt and the "shadow" infrastructure organization itself festered in to some of the awful behavior which ended up exploding in the saga with Susan et al. We promoted two senior engineers on my team in to management who proceeded to compete to be the first manager to hire a woman in to our 100 man, 0 woman engineering group. A year and a some change later Susan ended up reporting to one of them when we had all finally made up under the new SRE leadership.
I have some very surreal memories of these times, some dear memories of the brilliant folks I worked with, some unspeakably regretful memories of both, and some of this I had completely blocked out. It's hard to believe in the two year arc of this blog post, we hired endlessly (occasionally poorly, often brilliantly), we brought up four data centers including a reverse-engineering of our tech stack to launch two from first-principles in China, rebuilt entire parts of the company's architecture from empty git repos to a thousand plus servers in a quarter, and nearly burst at the seams doing it.
It's an irreplicable and irreplaceable experience but I can't say I'd go through with it again. I'm not even sure I'd wish it for someone: after three years on these infrastructure teams and two more on the privacy team after #deleteuber and starting work on GDPR compliance 6 months before the enforcement started, I left. two years later, I am still recovering from burnout spending down my little dragon's hoard healing a complete lack of interest in technical work.
I'm glad that Will has been distilling and expressing the lessons we learned so that others can go through this sort of experience in a less extreme and painful fashion.
More broadly, I think the level of granularity at which it makes sense to define terms (Mesos, Kubernetes, even Uber) has to roughly match the level of familiarity of the reader that will get something meaningful from the piece.
He's the CTO of a "wellness" tech company now and writes broadly on the topic of engineering and engineering management. The folks who will read this likely are familiar enough with these technical terms to glean knowledge, or at least to gloss over them.
Sad :(.
> That’s not to say the experience was entirely good. Working at that era of Uber extracted a toll. Many others paid a much worse price; I paid one too. I was in way over my head as a leader, and I struggled with it. On a work trip ramping up the Uber Lithuania office, my partner of seven years called to let me know they’d moved out. Writing had been my favorite hobby, but I gave up on writing during my time there. Some chapters enrich our lives without being something we’d repeat, and Uber is certainly one of those chapters for me.
As someone very early in their career, I would still take that experience in a heartbeat. Just by observation, I strongly believe that kind of experience is both extremely rare and extremely valuable.
I work for CloudKitchens, new startup of Travis Kalanick.
How you can write an entire article with an acronym in the title and not define it anywhere is beyond me.
You can adjust the strictness for audience and intent perhaps, but it's still sloppy not to.
This article could have helped me with that but it didn’t fully. It sounded more like a rant that will totally make sense maybe to 50% of infra or devops people.
Maybe it’s me though. I actually still don’t know what a true product manager is supposed to do lol.
While it's common practice for SRE teams to share some of the operational burden, it's for the purpose to find risks and vulnerabilities and then engineer solutions to those. I have never seen SRE do support.
SREs should know how to build and run reliable and maintainable systems. Architecture and design are the best tools for that.
This really only happens well at Google AFAICT.
the "new SRE" team Will mentioned was a bunch of ex-Google Infra folks who brought a holy book [https://sre.google/sre-book/table-of-contents/] with them and asked us to institute these practices whole-cloth. some of them had already been phased out of google by then but how could we have known
There's a lot of techniques and tools that are no longer in favor at Google not because they're not the best available SRE tools, but because that's how far Google has fallen from grace.
https://web.devopstopologies.com/
Note it shows anti types and useful types.
[1]: https://netflixtechblog.com/full-cycle-developers-at-netflix...
Managing and scaling "software infrastructure" (Kubernetes, cloud services, databases/caches, job/message brokers, etc). SREs tend to have expertise in these systems that generalist SWEs do not.
Incident response. Many production incidents are not caused by a bug in application code, but due to some cascading failure of some piece of the software infrastructure, often due to unexpected load, performance regressions or network/hardware outages. Because SREs have a broader picture of how the various pieces of infra fit together, they're best suited to start root cause analysis and determine whether it's an infra/code issue or some combination of the two.
Devops/developer productivity. Many SREs work on build/release systems, internal tools and enforcing best practices.
Perhaps it's only common within silicon valley circles, that being said, the fact that the author didn't define the acronym didn't help either.
I'm surprised you've never come across the term SRE/Production Engineer before (Assuming you meant to say "have only __just__ heard about what an SRE is"). I live in New Zealand and I'm pretty sure I knew what an SRE was before I even joined a big tech company (Not a SV one btw).
Edit: It's probably been around longer that SDET!
That the title wouldn’t be well known in Europe, or fully known in the US isn’t surprising to me. I’d probably expect someone who works in the start up world to know it, and would expect anyone working at the FANG compensation level to know it, but I don’t think it’s super common among the vast middle of companies.
(And for reference, I had know idea what SDET meant until I looked it up right now, although I certainly have worked with Software Development Engineers in Test. Probably an even more obscure term?)
I guess that's fair.
> (And for reference, I had know idea what SDET meant until I looked it up right now, although I certainly have worked with Software Development Engineers in Test. Probably an even more obscure term?)
I'm not sure - I work as an SDET and have found work as an SDET in Oxford and London England, remote in other European countries, and am now interviewing for several companies in the Bay Area. It seems pretty universal to me! But again, maybe that's just down to the bubble one is in. FWIW, it just means a software developer who focuses on testing as opposed to frontend or backend, so managing test infrastructure, reporting, writing frameworks for devs to use for testing, etc etc...
Wait! No. My vote goes to AI Ethicist.
At one point in my career I moved from a technical presales role to SRE. The engineers congratulated me on the promotion, while the sales people all thought it was a demotion. (And in terms of career progression it probably was.)
I'd love to see some evidence of that. I have never been in an organization where Ops folks and SREs are treated as if those roles are prestigious by other people in the engineering org. Generally they're treated like digital janitors, and assumed to be bad at software development. I've literally had an SWE /at work/ say to my face that "You're in Ops because you weren't smart enough to pass a code interview" because they were upset I had pointed out a flaw in their design and were essentially making an appeal to the nebulous authority of their title.
SRE is definitely more prestigious than being in technical presales to other engineers, but that's because to other engineers any role in the sales org is not real engineering (even if it is), and to sales people anything in the engineering org is a demotion because it means you are joining "the basement people". Sales people aspire to becoming a CEO, not to building something that makes a mark on the world. Engineers want to build the thing, they don't care /how/ they have to go about doing it (whether that's being a founder/ceo or working in the basement).
I get increasingly irritated with the myth that 'DevOps' are any different than the sysadmins of 10 years ago.
"But DevOps can code", yes, so could sysadmins, in fact, terraform, ansible, vagrant, saltstack, chef, puppet etc;etc;etc are all made by people who held the title of sysadmin when they were written.
In fact even the term "DevOps" was originally from a conference, where the idea was that "we can do systems administration in an agile way" -- NOTHING to do with coding, everything to do with getting developers and sysadmins working closely together in an iterative fashion.
I would personally be very happy being called a sysadmin, but doing so is career suicide, because we as an industry have decided that sysadmins are somehow braindead, and that you really need "SREs" or "DevOps" -- despite the fact that these are the same people.
What gets my goat even more is that people hate on sysadmins because of corporate culture, echos of centralised IT organisations that said no to everything.
But we're doing exactly the same thing with these new titles now. It's a joke.
(This text, from yesterday: https://news.ycombinator.com/item?id=31194067 )
——-
I would be very happy to be called a sysadmin again, we can take some principles from the SRE book and that’s fine.
This is increasingly annoying because oh can’t be called a “engineer” in Canada, it’s a protected title.
I even wrote a ranty blog post about this a couple years ago, I will post it here later.
You can’t offer engineering services to the public without being licensed.
Non-PEO licensed software engineers exist in large quantities in Canada without any issues. They provide services to a company, not the public.
I don’t think it’s solely a Unix/windows thing, but I still work with windows admins who manually set up machines. Like they set up a vm in azure and then spend hours or days pushing buttons until it’s “ready.” For one vm.
A person like this should not be in ops. They should be trained and improved to automate this by code.
So the world is full of “there’s always been devops” because 30 years ago sysadmins were scripting out their stuff in shell scripts and whatnot. And also “we’re moving to devops” where people still manually do system tasks.
I've been to meetups where I met staff from devops teams who talked about wanting to learn to code in a language like Python someday, to switch into a developer role, and whose response to me telling them I was a C developer was "wow that's really hardcore isn't it".
So perhaps there isn't an exceptation that all ops people can code, and perhaps those particular devops teams are not much different from the sysadmins of old.
But the term seems to have been diluted into almost meaninglessness as people seem to mean very different things when they say it.
SRE is related but my take is it’s an evolution/reaction to DevOps. Some took devops to mean “developers do all the the ops” and SRE puts on-call back onto a dedicated role (so back to the NOC and old school Ops) but keeping (and perhaps even further-emphasizing) the commitment to engineering work in service of automating operations at scale.
People definitely misuse labels. But I think that devops teams should be people who code.
Because SRE came out of Ops, which came out of SysAdmin, there's a cultural expectation both by other engineers and by management that you are essentially a digital janitor, and if you can code it's only minimally so that you can write glue scripts, rather than being equally capable as an SWE. The reality is that most of my peers were /better/ developers than equal-level SWEs, because they had a stronger understanding of lower level systems and their interactions than SWEs who spent all their time working in abstractions.
As another commenter notes, I would have been perfectly happy keeping the "SysAdmin" title my entire career, and while I held the title for awhile, it always grated on me when people were hired as DevOps Engineers (I would often tell managers "DevOps is a philosophy, not a job title."). At least SRE accurately describes the work in some abstract sense, you are an engineer and your primary duty is the reliability of the public facing site and its dependent backend services.
Frankly, the amount of derision I see here and in other tech circles from SWEs against basically every other role in tech makes me think SWEs are the ones with the ego, for that matter.
I can believe this. It’s especially funny when building a reliable system today requires a really basic understanding of few low level system concepts (eg graceful termination, signal handling, healthchecks, DNS) and most SWEs I work with simply refuse to learn these things.
Are the SWEs over worked with unrealistic deadlines? Also how did the SRE present this information to them? I think both factors can drastically have an effect on the outcome.
For example a small team who is already on the brink of being overburdened getting tossed a few links on Slack that says "please make sure you know about Kubernetes graceful shutdown timeouts and signals" with no other context is going to be met with resistance because now it feels like you're giving them another job. If they have no prior experience with that sort of thing it could take days to partially understand it. That's eating up time for their sprint, etc..
Alternatively, you could write a nice bit of documentation in private explaining the problem and why graceful shutdowns are a good thing, how it aligns with being able to deploy new versions of our apps, how they can create a better user experience for the apps they're building along with prepare anything you can to get as much information about your apps as possible (longest page responses, etc.), then spend a bit of time with your tech lead or someone who knows your applications well to iron out good values for all of your services (this assumes a worst case scenario where you don't have these metrics logged anywhere yet). Then wrap things up with a 5-10 min show / tell to share this knowledge with the devs so they have an awareness of it.
In my opinion that maximizes everyone's time, the SRE doesn't even need to know much at all about the app's business domain too since it all boils down to how long an app might take to exit. Knowing a bit about the tech stack can be researched for the document you'll write too, such as mentioning how you can hook into the shutdown process of the app server to potentially execute code during a graceful shutdown, etc..
I've found that after all of that, it was embraced in a very positive way with devs now looking for ways to improve the app to better handle the app shutting down at unknown times without data loss and minimal user disruption. This is now a skill set the devs can take with them anywhere around having experience building robust and resilient applications. Everyone wins (customers, devs, SREs, etc.).
The hypothesis appears to be spend money to bring in people who know how enterprises work. Instead of people “just” managing or managing managers, get “leaders” who are “proven”at managing business functions or lines business through managers of managers. So, pay big bucks to people that have risen to that size management before.
Of course this means “layering” the Wills of the world with LOB or functional managers who by virtue of having been at a big enterprise predating the unicorn thing, have no clue what the function needs to be (no ‘north star’).
Next thing you know, thanks to this overweening managerialism, the unicorn-turned-Leviathan is hiring “transformation” executives. You know what’s next… the next unicorn still made of Wills, ready to eat Leviathan’s lunch.
https://www.thestreet.com/investing/how-much-of-uber-does-sa...
> "In the interview with the digital news platform, Khosrowshahi said the 2019 murder of Washington Post journalist Jamal Khashoggi shows that "the (Saudi) government said they made a mistake. It's a serious mistake, but we've made serious mistakes, too right?"
Tech isn't some neutral zone free of ethical quandries. Google leadership pushing a Dragonfly contract with China is another example. Israel selling their Pegasus spyware to the Saudis so they can round up and murder dissidents is another example. Just because it's profitable doesn't mean that justifies doing it.