Uber reinvented a ton of technology because of the idea that off-the-shelf solutions wouldn’t work at “Uber scale”. However, there was very little accountability for whether the in-house solutions were necessary, and working on these tools would get you promoted. So for every “Uber scale” problem that a team actually solved, there were a couple other projects that were just half-baked alternatives to the off-the-shelf software that they should be using.
It turns out that “Uber scale” is not really that large, despite the name. But engineers kept repeating “Uber scale” and building infrastructure.
The same problem occurs at the larger tech companies like Google, Facebook, Microsoft, Amazon, and Apple, but in different ways and to different degrees. And to a large extent, engineers are copying what other companies do, and bringing ideas from one company to another when they hang out after work or switch jobs. For example, you can bet that these companies mostly have their own containerization and scheduling systems, many of which are undoubtedly not competitive with Docker or K8s in 2020, but K8s only goes back to 2014 and all these companies are older. I’m sure Borg and Tupperware are great if you work at Google or Facebook but I’m also sure that they’re missing a bunch of tooling that you’re used to. Same thing with build systems. Bazel, Buck, Pants, Please, and that Frankensteined system that Chrome uses are all copies of each other but Bazel is the only one with a decent size community and ecosystem, as far as I can tell.
Uber is absolutely not on the scale of those companies, and most of the time they should probably be using off-the-shelf solutions when they become available.
It's worth noting that none of these were (really) NIH. Blaze wasn't open source originally, and Buck, Pants, and Please were all essentially reimplementations of the closed source Blaze rebuilt by xooglers who went on to work at FB, Twitter, and presumably thought machine, though IDK if it's the same.
Then/concurrently, Google open sourced Bazel, which is mostly-blaze.
My personal sense is that this even applies to stuff like Borg, but who knows?
I've even written a few libraries like that, and still stuff out there isn't that good or does not exist years later, because it's too niche, yet still a generic issue in the industry. Can't open source it although.
In an alternative timeline, Buck for example could of been the OSS solution that google decided to devote their resources to, and then in the end it would of been a good idea that FB created their own version of it.
Mind you, Kubernetes was originally made written at Google by engineers that previously worked on Borg. So they will naturally design it to address problems with Borg. Kubernetes’s original project code name was “Seven of Nine”, because Seven of Nine is a friendlier Borg.
I particularly can't agree that k8s is a necessarily friendlier Borg. Borg is on rails; there are fewer choices for the user to make. K8s tries to be flexible enough to be used in various production environments, which leads to it being harder to use rather than easier.
Sometimes you also get issues on k8s where a google engineer responds "it's not done yet because we really, really don't want to repeat what a fiasco it was under Borg".
So I'd say it might be better even in direct comparison ;)
Google has enough project churn internally that services and programs do stand a good chance of getting replaced, though. There’s plenty of history.
If Google made a good build system with features not available in open source, eventually open-sourced it, continue to use it, and lots of other people are eager to use it, how does that demonstrate that they shouldn't have invented a build system?
“Problem with X” is not the same thing as “you should not do X”.
I’m also not getting why you think that I’m ignorant of Bazel’s origins, specifically, since I was talking about Borg.
Regarding Borg vs k8s, it's worth noting that k8s was never built to replace Borg or to even try to match the feature set of Borg. There are some nice things about k8s (like the ease of running an entire cluster on your local machine), but since k8s tries to be all things to all people it's not even close to being viable (or competitive with what Borg can do) for Google itself.
I agree, but I'd also add that the state of the world in 2016 when most of these infra projects were launched is very different from the state of the world in 2020. Back then, uber was not easily able to run on a bunch of OSS/cloud native solutions, but now for sure its not just possible but likely the best/most responsible way to do it [source: worked there for 4+ years]
Amazon EC2 was first offered in 2006 and was super new and immature for a number of years. Kubernetes was first released in 2014. Mesos was a research idea for a number of years starting in 2009 and didn't reach version 1 until mid 2016. Etc. Etc.
These companies generally invent their own solutions because the solutions everyone here thinks they should be using didn't exist or were not stable at the time they had a problem they needed to solve.
By the time those new solutions exist and are stable enough, it requires quite a bit of investment to migrate to the newer solution as these companies already have much of their tech stack stable and bringing in revenue on their in-house solutions.
In the latter case there’s still plenty of risk: you might be spending valuable time building something that doesn’t work, is obsoleted by future change or future products; and the one that most people dismiss, you will be at the mercy of its creators. The people who build your critical technologies will have power over the organization you might regret later, and if they do leave, finding replacements can be hard and/or expensive.
And once you have a whole system built, you're really going to have a hard time justifying re-doing your whole architecture to move to a product that does the same thing. Basically that only happens when you have a new head of engineering who comes in and goes, i don't care how many engineers you have to throw at it, and i don't care how much custom functionality you depend on that you lose, just figure how to replicate everything you do today using k8s.
That's not to discount the engineers working there, they indeed do have a number of bright engineers. But the no accountability comment is spot on.
I have not had the privilege of working for a "tech company" where tech and tech employees are first class citizens. I'm pretty certain the in-house tooling would be much better at a company like Uber or FAANG than a company like mine.
Everything was seamless - I'd log in, tap my security key, and instantly have access to pretty much the entire monorepo and the rest of the production + deploy systems. All of the internal tools integrated with each other -- for instance, I could create a CL (i.e. a pull request) and then fix issues raised by the CI system, entirely from the IDE.
"Owning the stack" completely in-house also extended to hardware -- all of my builds happened inside Google datacenters too (not on my machine), and the development box they provided me was a Chromebook.
It’s so good that it would actually be a factor if I were ever considering working somewhere else. Getting to work with Google’s in-house dev tooling is probably worth 15-20k to me. A lot more than 20k if the other company has a reputation for horrible tooling.
To fix the tooling? Maybe..
Granted, their DNA is still in hardware/devices/OS level stuff versus "services" despite all they claim to be, but how much more quickly could they move if they "owned more" of their stack?
That said, I'm also looking forward to "VSCode front-end in the browser" becoming ubiquitous. I remember using Visual Studio back in my first job in high school, and Intellisense made coding so much more... explorable and approachable.
Also worked at one that developed a ton of tools internally, but I wouldn't call many of them terrible.
- uber is a 50 billion dollar company and they have to de-risk themselves by owning entire stacks, top to bottom. if it means re-creating something from scratch... who cares, they have billions.
My take on it is that the engineers want to make complicated solutions to hard problems to justify their salaries and get promoted, and that managers encourage that behavior so they can defend their headcount. I’m not accusing any of these people of acting in bad faith here—nobody’s reinventing tech to sabotage the company, it’s just that the system encourages this kind of behavior.
This problem is not unique to Uber, you’ll see similar things happen across the industry to different degrees.
Not what I said. Let’s move on.
> That's a dangerous philosophy…
Go pick a fight with someone else.
I still shake my head whenever I hear about how Uber built an entire Slack clone for "Uber-scale".
As a company, you have to invest immense resources in evolving the infrastructure and it is a never-ending task: the more capacity you build, the more your organization will grow in ways that tax that capacity. The goal posts are ever-receding.
As an individual developer, all that fancy distributed infrastructure becomes a barrier that must be laboriously overcome at every step. Your expectations about what constitutes a "fast" operation get gradually distorted. Working on "small" side-projects (a few thousand files or less) feels effortless in comparison.
Not to say you shouldn't seek out the "big tech" experience; just that it is not all roses.
And a lot of tearing your hair out when trying to use tools in the ways that they weren't intended to be used.
Of all languages, Go is not the language where you want to be doing things against the grain (that is, against the way the language designers intend it to be used). Some people like that about Go, and some people don't. Either way, Go is a very opinionated language with a very opinionated ecosystem, and ignoring those is a recipe for frustration.
Last time I checked, Bazel was not recommended by the Go developers - and with good reason: there are a lot of gaps in rules_go/Gazelle, which this blog post alludes to but glosses over. While Google uses Blaze internally, Blaze is not Bazel, and the differences are very apparent to anyone who has used both to build Go specifically. Furthermore, rules_go was developed entirely independently of the Blaze ruleset that Google uses internally, so it's a tool that's not actually used internally at Google, but also not used by the majority of the non-Google Go developer community either, in addition to not being recommended by the Go team at Google[0].
[0] There is exactly one mention of Bazel on the entire golang.org domain - in a changelog from over two years ago, in an /x/ package, where Bazel is mentioned as one of two build systems in a "such as" clause that the new package could potentially enable support for (/x/ packages are considered experimental and not subject to the same backwards compatibility or maintenance guarantees as the rest of the project).
Also anecdotally, bazel and golang work really well together IME. The community seems pretty active, and the upsides of using gazelle/bazel with golang seem to outweigh any downsides (though I'd be hard pressed to name a downside, that isn't inherit to golang itself).
This is really not my experience from having used Bazel with Go for the last four years. But I'm happy you are are apparently not running into issues.
except when it comes to dependency management
We're actively hiring for a senior SWE right now, so feel free to shoot me a note if you're looking.
I worked at a company that used protos from day one for example.
Many large companies are still using JSON w/ schemas as their network serialization layer just fine.
One companies boring tool is another companies favourite hyped third-party tech tool. (Until it doesn't work for them)
Just like Kubernetes, Bazel is not a silver bullet. But if you get all your ducks in a row, switching over is a gamechanger and will help you out even at the smallest scale.
Having spent the last few weeks of my life converting our build system to Bazel I guarantee you it's far from what I would consider cool. It's more of a necessary evil because the alternatives are even worse.
FANG such as FB and Google has some big monorepos, and you can try that if you are interested.
How so, in 2020?