Building Uber’s Go Monorepo with Bazel
eng.uber.com
eng.uber.com
That changes drastically if your codebase contains a lot of C++ and your SW model doesn't quite match the way Bazel tries to push you into. For instance doing any of the following things will quickly turn into a nightmare when using the Bazel C++ rules:
* dynamically linking libraries
* using a containerized approach for library includes (so transitive relative include paths)
* using different toolchains and cross compiling
* interfacing with thirdparty libraries
We are talking seasoned build engineers ending up frustrated after literally months of trying to achieve something that is easy as pie in CMake.
In addition there is still no real IDE support. The CLion plugin is permanently broken and lags behind versions. No real VSCode support. Using custom rules makes this even worse, to the point where CLion will refuse to sync and not having a way to produce a compilation database json.
There are so many bugs open on those projects and no progress or answers. I can not recommend Bazel if C++ or C is what you care about.
But it's not the same as CMake. A properly set up Bazel project gives you so much more: organization-wide incremental builds, build cache and build farm. Also, actual full hermeticity of builds (no, taking ambient deps from a Docker container doesn't replace that).
Comparing CMake to Bazel is like comparing some barely-working bash scripts on a single box to a kubernetes deployment. Maybe you're okay with just bash scripts, but some of us aren't, and that's where Bazel comes in.
All the complex things that Bazel brings into the game (distributed builds and caches primarily) simply don't have any value when the main task that you're trying to make faster is not done well.
I would much rather like to see more tooling that does these things but builds upon layers of battle-tested existing software.
E.g. Google's protobuf libraries [1] can be built with Bazel, and they depend heavily on system headers outside the repository, e.g. <iostream> and <stdio.h> and lots more. If those headers subsequently change, Bazel will not pick up on this and will not know to rebuild the parts of the build that depend on them.
To reproduce: run `bazel build -c opt //:protoc_lib` and then put random garbage in your /usr/include/stdio.h and /usr/include/c++/<version>/iostream and then rerun the bazel command -- it will not know to invalidate the build cache. If you `bazel clean` and then build again, you'll get different results.
Bazel does a lot of really nice things and I can believe that within a google3-like environment (where the source code never references a system header?), it effectively provides hermetic builds, but in practice as used outside Google (or even in Google's public OSS releases) it doesn't seem to really match this description or enforce a hermetic seal. What am I missing?
In my view the challenge here is that a dependency changes (e.g. /usr/include/stdio.h is upgraded by the system package manager, or two users sharing a cache have different versions of a system library) and Bazel doesn't realize that it needs to rebuild. It would be a pretty heavy hammer if the way to fix that requires every OSS project to include the whole user-space OS (C++ compiler, system headers, libraries) in the repo or via submodule and then be careful that no include path or any part of the build system accidentally references any header or library outside the repository.
And maybe this issue just doesn't need to be fixed (it's not like automake produces build rules that explicitly depend on system headers either!) -- my quibble was with the notion that Bazel, unlike CMake or whatever, provides fully hermetic builds, or tracks dependencies carefully enough to provide an organization-wide build cache across diversely configured/upgraded systems.
No, that's up to the user. If you include protobuf in your Bazel workspace and have a toolchain configured, protobuf will be built with that toolchain (this is also how you would cross-compile). Bazel workspaces are still a little rough around the edges and interactions between them can be very confusing, as Google doesn't need them internally (everything is vendored as part of the same tree).
Under the hood there's a default auto-configured toolchain that finds whatever is installed locally in the system. Since it has no way of knowing what files an arbitrary "cc" might depend on, you lose hermeticity by using it.
Of course, the whole model falls apart if you want to depend on things outside of Bazel's control. However in practice I've found that writing Bazel build files for third-party dependencies isn't as hard as it seems, and provides tons of benefits as parent mentioned.
For example, last week I was tracking down memory corruption in an external dependency. With a simple `bazel run --config asan -c dbg //foo`, I had foo and all of its transitive dependencies re-compiled and running under AddressSanitizer within minutes. How long would that take with other build systems? Probably a solid week.
I think the reason the public finds this so mysterious is the documentation for CROSSTOOL is terrible, and virtually all of the people who learned how to use blaze at Google go out in the industry with no understanding at all of CROSSTOOL, because there's a small dedicated team who maintain it.
Doing it on a company scale is probably harder. Last place I worked used bazel but did not bother with remote.
Starting a new project with it, so it can help you do things the "right way" is a much more attractive proposition.
That said, the IDE/tooling situation for Bazel was disappointing to me in the Java world too. It definitely lags and is incomplete.
And while Nix has an insane learning curve, figuring out how to extend Bazel (and consequently build one's own Python plugin) was a nightmare. The documentation was unhelpful/incomplete and the codebase was incomprehensible and poorly organized, with inheritance sprinkled on everything for its own sake (which is to say, par for the course for Java apps).
I know... "Don't look a gift horse in the mouth! / Don't complain about open source software!" Fair enough, but I don't have to use it, and I'm free to caution others not to as well.
Unfortunately Kubenix was too immature, my knowledge of Nix insufficient, and Hydra required privileged containers due to Nix build sandboxing.
Have you used nix as a primary build system rather than a meta build system? If it works, I might have the energy to revive my project.
1. Set up sccache backing server (I use docker redis) 2. Export sccache config. Set SCCACHE_REDIS to a Redis url in format redis://[:<passwd>@]<hostname>[:port][/<db>]
Set CMAKE_<LANG>_COMPILER_LAUNCHER (cc, cxx) to sccache.
I think this is something that's slowly, but surely, changing. There's a lot of adoption of Bazel from some big players out there (Uber, Bloomberg, etc). The Blaze team @ Google has also started incorporating outside OSS maintainers into their OSS software (specifically rules_python).
Things are improving. Many of the things you've addressed (third party libraries, cross compiling, dynamically linking things) have been talked about by many people in Bazel's slack. It might not be flawless currently but Bazel's abstractions of how builds work are cool, RBE/caching is amazing for large projects, and a single build system with the promise of gracefully handling "every" language is something our field has been needing for some time.
In 3 to 10 years I can imagine being able to start writing code and never think about which language some library I want to use is written in, which package manager and toolchain I need to use to build everything, how to package my code to be deployable in production, how to setup unit tests, how to get autocompltions and docs and everything in my IDE, etc. The answer could just be "bazel, a language server, an IDE that speaks lsp".
We've been using Bazel not just to build and test our Go apps, but to build docker images, compile proto files, even deploy to k8s. It's really versatile and the developers only need to know one tool to build and test the whole environment.
Haven't run a bazel clean in... 6 months maybe?
The multi-language really is a killer feature when the project inevitably becomes polyglot, such as when doing protobuf/gRPC or using CGo.
As the blog post notes, Makefiles are the usual weapon of choice for Go projects of any sufficient complexity, but these are quickly outgrown.
Of course Bazel is far from perfect, in fact sometimes you may be better off running your own rules instead of the official ones in some cases, but IMO it makes a pretty decent build system for putting all of your code in one place.
On the other hand, it’s one of those Google things where if you haven’t seen it in action in a functional configuration it’s hard to explain why it’s actually nice. And sadly I feel unsatisfied with the Node.JS rules for example, which is probably how a lot of people will first experience Bazel (since I believe Angular supports Bazel this way.)
Bazel also come with a query tool that makes it easy to tell what other targets your package depends on. Go also has such a tool, but again it's limited to Go dependencies.
I believe Uber has been using bazel for awhile now, and probably were invested in it well before the Go module cache was a thing, so Bazel's build cache is an important optimization.
That said, I do agree with other posters that Bazel + go feels like nobody's priority. I use gopls and love it very much, but it doesn't understand how to ask Bazel where files are to analyze. A bug has been open for years, lots of people saying "I'll put it on my OKRs this quarter", and it's just not getting done. I am close to just fixing it myself.
But I have a feeling that this happens to everyone. I have worked on other Bazel projects and they all have detailed setup instructions; for example, Envoy uses Bazel but somehow shells out to cmake, so you have to have the right version of cmake, the right version of clang, etc. to produce a build. To me that defeats the purpose of using Bazel; Bazel should build the toolchain and all dependencies for you, otherwise what's the point? All you should need is Bazel and the source code. But that is not how people are using it. I don't know what the canonical successful Bazel project is, but I haven't found one yet.
I am currently working on a Typescript/Go project with gRPC/Web and will see if Bazel gets me that "one command full build" dream. I am somewhat motivated to make it work, but I have a feeling that there will be warts that makes me regret it. I used Bazel (well, the internal version) when I was at Google and every time I used it it was a joy. Incremental builds where you edited one file and wanted to run the tests were very fast, and builds where you modified some core library and wanted to see the impact across all the affected projects were also fast (easy to use 100s of CPU hours in just a few minutes of wall clock time).
Ultimately I think Bazel is the right approach, but I think to be successful you have to commit to having one engineer work on tooling full time. At your startup, probably not feasible. At Uber, probably very feasible.
(Going off topic, I would love to see what other people are doing for polyglot projects, because I feel like everything is a polyglot project these days. I have never seen a setup I liked -- install one program, check out the code, have a guaranteed working build that is bit-identical to the production image. The fact that nobody does this scares me deeply. But that's where we seem to be as a field.
I am also interested in seeing people's templates for projects. For example, I would like to do a C++ microcontroller project, and want to see how people are building those. I need clangd support, I need to use libraries, and everything I've seen is "hack together a makefile and cross your fingers". It just scares me.)
And a lot of tearing your hair out when trying to use tools in the ways that they weren't intended to be used.
Of all languages, Go is not the language where you want to be doing things against the grain (that is, against the way the language designers intend it to be used). Some people like that about Go, and some people don't. Either way, Go is a very opinionated language with a very opinionated ecosystem, and ignoring those is a recipe for frustration.
Last time I checked, Bazel was not recommended by the Go developers - and with good reason: there are a lot of gaps in rules_go/Gazelle, which this blog post alludes to but glosses over. While Google uses Blaze internally, Blaze is not Bazel, and the differences are very apparent to anyone who has used both to build Go specifically. Furthermore, rules_go was developed entirely independently of the Blaze ruleset that Google uses internally, so it's a tool that's not actually used internally at Google, but also not used by the majority of the non-Google Go developer community either, in addition to not being recommended by the Go team at Google[0].
[0] There is exactly one mention of Bazel on the entire golang.org domain - in a changelog from over two years ago, in an /x/ package, where Bazel is mentioned as one of two build systems in a "such as" clause that the new package could potentially enable support for (/x/ packages are considered experimental and not subject to the same backwards compatibility or maintenance guarantees as the rest of the project).
Also anecdotally, bazel and golang work really well together IME. The community seems pretty active, and the upsides of using gazelle/bazel with golang seem to outweigh any downsides (though I'd be hard pressed to name a downside, that isn't inherit to golang itself).
This is really not my experience from having used Bazel with Go for the last four years. But I'm happy you are are apparently not running into issues.
except when it comes to dependency management
We're actively hiring for a senior SWE right now, so feel free to shoot me a note if you're looking.
I worked at a company that used protos from day one for example.
Many large companies are still using JSON w/ schemas as their network serialization layer just fine.
How so, in 2020?
Uber reinvented a ton of technology because of the idea that off-the-shelf solutions wouldn’t work at “Uber scale”. However, there was very little accountability for whether the in-house solutions were necessary, and working on these tools would get you promoted. So for every “Uber scale” problem that a team actually solved, there were a couple other projects that were just half-baked alternatives to the off-the-shelf software that they should be using.
It turns out that “Uber scale” is not really that large, despite the name. But engineers kept repeating “Uber scale” and building infrastructure.
The same problem occurs at the larger tech companies like Google, Facebook, Microsoft, Amazon, and Apple, but in different ways and to different degrees. And to a large extent, engineers are copying what other companies do, and bringing ideas from one company to another when they hang out after work or switch jobs. For example, you can bet that these companies mostly have their own containerization and scheduling systems, many of which are undoubtedly not competitive with Docker or K8s in 2020, but K8s only goes back to 2014 and all these companies are older. I’m sure Borg and Tupperware are great if you work at Google or Facebook but I’m also sure that they’re missing a bunch of tooling that you’re used to. Same thing with build systems. Bazel, Buck, Pants, Please, and that Frankensteined system that Chrome uses are all copies of each other but Bazel is the only one with a decent size community and ecosystem, as far as I can tell.
Uber is absolutely not on the scale of those companies, and most of the time they should probably be using off-the-shelf solutions when they become available.
- uber is a 50 billion dollar company and they have to de-risk themselves by owning entire stacks, top to bottom. if it means re-creating something from scratch... who cares, they have billions.
I still shake my head whenever I hear about how Uber built an entire Slack clone for "Uber-scale".
My take on it is that the engineers want to make complicated solutions to hard problems to justify their salaries and get promoted, and that managers encourage that behavior so they can defend their headcount. I’m not accusing any of these people of acting in bad faith here—nobody’s reinventing tech to sabotage the company, it’s just that the system encourages this kind of behavior.
This problem is not unique to Uber, you’ll see similar things happen across the industry to different degrees.
Not what I said. Let’s move on.
> That's a dangerous philosophy…
Go pick a fight with someone else.
I have not had the privilege of working for a "tech company" where tech and tech employees are first class citizens. I'm pretty certain the in-house tooling would be much better at a company like Uber or FAANG than a company like mine.
Everything was seamless - I'd log in, tap my security key, and instantly have access to pretty much the entire monorepo and the rest of the production + deploy systems. All of the internal tools integrated with each other -- for instance, I could create a CL (i.e. a pull request) and then fix issues raised by the CI system, entirely from the IDE.
"Owning the stack" completely in-house also extended to hardware -- all of my builds happened inside Google datacenters too (not on my machine), and the development box they provided me was a Chromebook.
Granted, their DNA is still in hardware/devices/OS level stuff versus "services" despite all they claim to be, but how much more quickly could they move if they "owned more" of their stack?
That said, I'm also looking forward to "VSCode front-end in the browser" becoming ubiquitous. I remember using Visual Studio back in my first job in high school, and Intellisense made coding so much more... explorable and approachable.
It’s so good that it would actually be a factor if I were ever considering working somewhere else. Getting to work with Google’s in-house dev tooling is probably worth 15-20k to me. A lot more than 20k if the other company has a reputation for horrible tooling.
To fix the tooling? Maybe..
Also worked at one that developed a ton of tools internally, but I wouldn't call many of them terrible.
Amazon EC2 was first offered in 2006 and was super new and immature for a number of years. Kubernetes was first released in 2014. Mesos was a research idea for a number of years starting in 2009 and didn't reach version 1 until mid 2016. Etc. Etc.
These companies generally invent their own solutions because the solutions everyone here thinks they should be using didn't exist or were not stable at the time they had a problem they needed to solve.
By the time those new solutions exist and are stable enough, it requires quite a bit of investment to migrate to the newer solution as these companies already have much of their tech stack stable and bringing in revenue on their in-house solutions.
In the latter case there’s still plenty of risk: you might be spending valuable time building something that doesn’t work, is obsoleted by future change or future products; and the one that most people dismiss, you will be at the mercy of its creators. The people who build your critical technologies will have power over the organization you might regret later, and if they do leave, finding replacements can be hard and/or expensive.
It's worth noting that none of these were (really) NIH. Blaze wasn't open source originally, and Buck, Pants, and Please were all essentially reimplementations of the closed source Blaze rebuilt by xooglers who went on to work at FB, Twitter, and presumably thought machine, though IDK if it's the same.
Then/concurrently, Google open sourced Bazel, which is mostly-blaze.
My personal sense is that this even applies to stuff like Borg, but who knows?
Mind you, Kubernetes was originally made written at Google by engineers that previously worked on Borg. So they will naturally design it to address problems with Borg. Kubernetes’s original project code name was “Seven of Nine”, because Seven of Nine is a friendlier Borg.
I particularly can't agree that k8s is a necessarily friendlier Borg. Borg is on rails; there are fewer choices for the user to make. K8s tries to be flexible enough to be used in various production environments, which leads to it being harder to use rather than easier.
Google has enough project churn internally that services and programs do stand a good chance of getting replaced, though. There’s plenty of history.
Sometimes you also get issues on k8s where a google engineer responds "it's not done yet because we really, really don't want to repeat what a fiasco it was under Borg".
So I'd say it might be better even in direct comparison ;)
I've even written a few libraries like that, and still stuff out there isn't that good or does not exist years later, because it's too niche, yet still a generic issue in the industry. Can't open source it although.
In an alternative timeline, Buck for example could of been the OSS solution that google decided to devote their resources to, and then in the end it would of been a good idea that FB created their own version of it.
If Google made a good build system with features not available in open source, eventually open-sourced it, continue to use it, and lots of other people are eager to use it, how does that demonstrate that they shouldn't have invented a build system?
“Problem with X” is not the same thing as “you should not do X”.
I’m also not getting why you think that I’m ignorant of Bazel’s origins, specifically, since I was talking about Borg.
Regarding Borg vs k8s, it's worth noting that k8s was never built to replace Borg or to even try to match the feature set of Borg. There are some nice things about k8s (like the ease of running an entire cluster on your local machine), but since k8s tries to be all things to all people it's not even close to being viable (or competitive with what Borg can do) for Google itself.
I agree, but I'd also add that the state of the world in 2016 when most of these infra projects were launched is very different from the state of the world in 2020. Back then, uber was not easily able to run on a bunch of OSS/cloud native solutions, but now for sure its not just possible but likely the best/most responsible way to do it [source: worked there for 4+ years]
And once you have a whole system built, you're really going to have a hard time justifying re-doing your whole architecture to move to a product that does the same thing. Basically that only happens when you have a new head of engineering who comes in and goes, i don't care how many engineers you have to throw at it, and i don't care how much custom functionality you depend on that you lose, just figure how to replicate everything you do today using k8s.
That's not to discount the engineers working there, they indeed do have a number of bright engineers. But the no accountability comment is spot on.
One companies boring tool is another companies favourite hyped third-party tech tool. (Until it doesn't work for them)
As a company, you have to invest immense resources in evolving the infrastructure and it is a never-ending task: the more capacity you build, the more your organization will grow in ways that tax that capacity. The goal posts are ever-receding.
As an individual developer, all that fancy distributed infrastructure becomes a barrier that must be laboriously overcome at every step. Your expectations about what constitutes a "fast" operation get gradually distorted. Working on "small" side-projects (a few thousand files or less) feels effortless in comparison.
Not to say you shouldn't seek out the "big tech" experience; just that it is not all roses.
FANG such as FB and Google has some big monorepos, and you can try that if you are interested.
Just like Kubernetes, Bazel is not a silver bullet. But if you get all your ducks in a row, switching over is a gamechanger and will help you out even at the smallest scale.
Having spent the last few weeks of my life converting our build system to Bazel I guarantee you it's far from what I would consider cool. It's more of a necessary evil because the alternatives are even worse.
That's a weird way to put it. Bazel is the open source variant of what Google uses for its monorepo, which is several orders of magnitude larger and arguably more complex.
That Google had to come up with Bazel for themselves only supports that point. It isn't obvious whether it'd then be applicable to other existing cases without customization
Lots of knobs that vary here tho, such as where one draws the line between configuration & customization
https://docs.bazel.build/versions/3.1.0/remote-execution.htm... https://docs.bazel.build/versions/3.1.0/remote-caching.html
What's not open-sourced is mostly the interaction with Google internal infrastructure.
I have been tinkering with a few ideas to make the existing tooling work with Bazel, but the effort is larger than I had originally expected.
https://github.com/earthly/earthly
Disclaimer: I am Earthly's creator.
I work on a project with a Go backend and a React frontend, and having to update all those React dependencies myself is what keeps me from moving to a system like Bazel.
I have not seen metris or blogs that prove this "uptick in build efficiency" or an increase productivity.
While I do like the idea behind bazel, I hate repeating deps in things like "go_repository" with gazelle.
No go_repository duplication.
Monorepos are not efficient. They are easier to manage when a team is small but as the team grows and you have more and more deliverables with separate versioning you are introducing control structures in your automation. Complexity explodes!
Anyway, all this does matter if you don't make any profit :)
I like to look at Google's GitHub commit messages to get an idea of the pace of their revision history. Yesterday they committed something with a Piper revision of 311324901. A month ago it was 306514102, and a year ago it was 248381230. That's about 160k revision numbers per day.
Disclaimer: ex-googler.
If you need to build N libraries and M executables then you are going to have N+M targets (assuming building for one arch only) whether you use a monorepo or not... Either way you are going to issue N+M build commands.
Also, in bazel you can do
bazel build //...
to do a full build of all targets in a workspace. If they are not doing this but instead passing each target name individually then that probably means that they are only building a subset of everything, and even that subset is too much to pass in a single command line invocation. I'll grant you that it seems excessive to have that many targets, but again I don't see having so many targets as an explicit issue of the monorepo.The complexity exists whether it's a monorepo or many separate repos. A monorepo lets you encode that complexity as versioned code in your build system. Separate repos encode it across people's heads, wikis and who knows what else. Hiding complexity doesn't mean it doesn't exist, just that it will bite you 10 times as hard eventually.