Scaling Rust Builds with Bazel
mmapped.blog
mmapped.blog
Maybe you have some really complicated use case that makes cargo insufficient, but really, it's pretty damn nice and it's integrated with everything.
There are languages that need bazel, rust isn't one of them
Though I seriously wish there were a better option, because there are a _lot_ of usability things with Bazel that make me very frustrated when working with a repo that is "merely" quite large.
It still doesn't support post build steps, and integrates poorly with other tools that are used for more general builds.
It's just, when I joined Google I found a build system I didn't hate. This was the first time. I don't actually love Bazel (Blaze internally), but I love that I don't hate it.
Every time I have to learn a new build system in open source work I groan. Why are there so many systems for essentially the same thing?
Only lcov is supported, and due to the sandboxing it's impossible to get other tools working since they always spit out extra instrumentation data that's neither cached nor available on the next sandbox.
Does blaze support code coverage properly? Does Google just not care about code coverage? Am I just missing some obvious thing in the documentation?
Code coverage is pretty well supported in the internal ecosystem but I would not be surprised to hear this is because we designed Bazel around the system Google was already using for coverage.
What you're actually saying is that internally Google managed to get a single and very specific use case to work and forced all projects to either adopt it or have no code coverage.
I forget the exact details but I think you can do this with a custom run_under test wrapper used during coverage. The wrapper runs your test, with whatever custom coverage system you have, then afterwards it converts to LCOV. Sometimes this might require updates to language specific rules because the test runner itself needs to be configured.
The bottom line is that Bazel upstream does not put a lot of attention into this, so it’s under developed outside of Google, because Google internal does not use the OSS rules.
The testing is not all you need to modify. You need to modify every compile and link step, since you get byproducts of the instrumentation on every step.
>Sometimes this might require updates to language specific rules because the test runner itself needs to be configured.
That's what I found, too, however I feel like rewriting the C++ rules is a huge undertaking, since they're not defined in Starlark.
I have to call bullshit on this. With C or C++ you need to build the whole project with specific compiler flags to be able to generate code coverage reports. This is not something you implement by tweaking a setting in unit tests.
The most useful leads I've found so far are:
* The design doc for the (abandoned?) coverage framework [1]
* A prototype of GCOVR support [2]
* Comment on getting llvm-profdata working [3]
[1] https://docs.google.com/document/d/1-ZWHF-Q-qCKf19ik-t33ie58...
[2] https://github.com/ulfjack/bazel/commit/e9f21bbf562c1f6006eb...
[3] https://github.com/bazelbuild/bazel/issues/8178#issuecomment...
There is definitely a need for better (and Meson is the foss answer, as well as Shake). But Bazel need a context that OpenSource do not have. Time and a clean slate of dependencies cooperating.
Also noone fund tools for DX for OpenSource devs :) only for JS and even that...
But the open source world has the resources to write new build systems (and ecosystems!) time and time again. This argument does not hold up.
What leads you to believe that writing new build systems is a popular hobby?
I mean, cmake is the de facto standard for C++ projects and it has been for what? Over a decade? What does this tell you?
It’s not solely the user’s responsibility to hold it right. But Bazel is and has been developed in the “open source context” for long enough now that your argument doesn’t hold up imo.
Tools commonly available in distros are typically much easier to use for outsiders.
Definitely a ton of rough edges on Bazel but the passage of time is rounding them, even if progress uneven and frustratingly slow.
> Tools commonly available in distros are typically much easier to use for outsiders.
Which part about autotools or ninja is easier?
There are some new talks coming out external from the Bazel maintainers which do a better job of this, but you still have to hunt for them.
This is made worse because each of the rules packages, which are super useful, have their own idiosyncrasies which only make sense with a solid foundation in what problems Bazel solves, how it works and how to write/read rules.
My other wishlist for Bazel was an easier time setting up a hermetic sandbox without having caveats. This is improving, so I am optimistic on this front.
For context, I am not a Google (or ex-Google) engineer and I’ve had to work with a ton of different build systems on different levels of the stack.
We had a googler at our company switch everything to bazel and it just slowed everyone down.
There really is something about Google engineers, this sort of subtle arrogance that the Google way is superior to everything else is existence. Maybe I'm the ass hole but I wonder if anyone can relate?
I'd love to hear if anyone can make a case for Bazel knowing that cmake+build cache tools like ccache buy far better performance improvements than any full or incremental build.
It's just that I really really hate maintaining Makefiles and CMake configs.
Admittedly in the latter case it's because I have never grokked CMake, but also... I don't want to grok CMake, I grok a perfectly good way of building C++ already. And it can also build Go and Rust and Typescript and Proto and JSonnet and it can run my janky codegen shell scripts and if we migrate to Carbon it will build that too and I won't even have to read the documentation.
Not saying I think everyone should migrate to Bazel (I hope I will not become That Ex-Googler), just explaining where the urge comes from.
CMake is a high level makefile/build system generator.
There is nothing to maintain in modern cmake. You just set what executables/libraries you want to build, set their dependencies, and you're done.
And of all alternatives you chose to push, you pick Bazel of all things? Feature-incomplete and requiring all sorts of low-level maintenance?
The logic is convoluted, documentation is over complicated and the language contains too many primitive functions and ways to shoot yourself in the foot.
There's really nothing good that's available.
When I use Bazel it is because:
- I want to build nearly everything from source in a controlled hermetic environment. No more "well it works on my machine"... only to learn the developer's libs leaked into the build environment.
- I want to do this with a single build system that works across many languages, and to build libraries which are used in many different languages.
- I want to do this within a sandboxed environment for peace of mind when compiling thousands of dependencies.
- I am willing to pay the money and time to setup remote executors to keep performance acceptable.
The issue is that to make it really nice:
- you have to convert everything to it, not just almost everything
- you need to set up build servers and lightning fast caching servers
- you need a team dedicated to building bazel rules for whatever doesn't already have rules (see: you must convert everything). This team is a bottleneck
Regular google engineers aren't exposed to these issues, because inside google it just works. It's only once they try to deploy it themselves that they realize it's a sisyphean task.
The thing is, bazel actually is really good at solving the issues it was designed for, it's just most companies do not have those particular problems acutely enough that it's worth paying the cost to switch to bazel.
This is a major issue. A good build tool shouldn't require a dedicated team.
Ideally a build tool should function as sort of a settings menu for the project.
Build systems are complex because they solve issues which only exist because of how OSes handle libraries and executables. I am pessimistic we can ever solve that issue in a capitalist society. I will be optimistic if we have a new major OS that isn't Windows, Linux, or MacOS in our lifetime or we solve death.
Why?
The world has been running on make for decades, and it runs just fine. Meanwhile, does Bazel support basic stuff like shared libs or vended libs that are not deployed system-wide? Last time I checked, it didn't.
Can you actually provide a working example?
I'd love to hear what leads you to believe that adding yet another build system in the form of Bazel would solve that problem for you.
I agree that Cargo is already near perfect and I wouldn't suggest you switch to Bazel if you have a Rust project.
In our case we have a C++ project with some Python, some Rust. Having everything in Bazel here is good for us.
It's a good thing then that cmake+ccache work out of the box by passing the caching tool of your choice like ccache to the compiler launchers.
Do you think that the problems with cargo mentioned in TFA are exaggerated or solvable? I think Bazel's support for caching, parallelism and custom rules are real advantages for large projects. What's wrong with this argument wrt Rust/cargo?
I'd say they're exaggerated at least slightly, yes.
It starts out by saying "Cargo isn't a build system" (which is already the sort of provocative exaggeration designed to increase participation), and then immediately veers off and starts talking about running tests. Cargo is a build system and dependency manager. It also happens to provide `cargo test` for running unit tests and integration tests, but this support is relatively simple, and if you need to do something weird then you're welcome to write a script to serve as your test runner. But a build system is not a test runner; a build system produces artifacts, and a test runner runs tests. The author appears to be annoyed that Cargo is not a build system for web assembly, which is part of their test process, and obviously Cargo is not. This is a somewhat eye-rolling critique, and does little to motivate rewriting the world in Bazel.
The second complaint is that Cargo uses file timestamps to decide when to recompile files. I suppose it is possible that Cargo could then proceed to lex any changed file to produce a comment-free token stream, hash the token stream, and then compare that hash to a previously-computed hash that you've stored in a database somewhere. Frankly, this sounds like a lot of work for essentially zero benefit; how often, exactly, do you go into a file and change only a comment? This is a weak justification.
It then mentions how changing git branches can invalidate the cache, but I don't think that's true. Git only updates timestamps on files that have changed when changing branches. If a file has changed when changing a branch, that is, again, probably because it has actual changes, and will almost never be because only a comment has changed.
It then makes some suggestions about Cargo not properly tracking changes to some global system resource that they have, which is simply too vague to comment on, and the fact that their prior reasons have been so weak makes me think that if this, of all things, is what they're trying to be vague about, then I assume this must be so unique and specific that even they recognize it's a weak critique.
Bazel is a nightmare. I wouldn't wish it on my worst enemy.
Don't look at headcount to gauge complexity, difficulty, etc.
What I've experienced is the larger the team the more likely to have dilution of ownership. The less ownership the less likely each person will spend the energy or fully understand solving root causes.
In two ways, one being that more people legit need a more complicated system. You want that so that they interact with the system, not with each other.
The other being a bit of a Parkinson's Law. That is specifically that you will expand work to fit the allocated time; but I posit that you will also expand work to fit the people doing it. So, more people pushing ideas into the codebase will keep more ideas in the codebase. Even if fewer would work.
What parts of Bazel were a nightmare? What problems were fixed by these 4-5 people that were previously impacting all 25 people routinely?
Hey look another nimwit that make stupid assumptions without understanding the nature of the problem being solved.
> and it's integrated with everything.
Oh wow Cargo integrates with Python, shell scripts, C++ and node JS? Holly sh*t that’s amazing.
/s
Speak for yourself. Our Rust monorepo has nearly a dozen apps and counting, lots of shared libraries, C++ dependencies, etc. Tests require multiple cargo invocations.
We need a proper build system with distributed build caching. Cargo takes forever and doesn't understand the workloads.
I feel like our use case at Grapl wasn't particularly crazy. It was just slightly more than "I'm a dev building a library and publishing it" - we had a workspace with multiple devs and CI/CD. CI/CD was really slow - issues with caching, rerunning unnecessary tests, etc.
Plus we had some Python code so unifying our builds would have been ncie.
To add onto this, Grapl had a moderate sized codebase with less than 10 engineers. I do think we could have done some optimizations in cargo, but it's simply much easier to use a tool like Bazel/pants etc that does the work for you as long as it supports your language of choice.
What problems, exactly? The author acknowledged cargo works but hand-waves over vague claims of "it does not track dependencies well or support arbitrary build graphs", followed by acknowledging those issues are either non-issues or solved problems whose solution was dismissed without any good reason (caching tools).
The author's claims boil down to "I wanted to shoehorn Bazel but had no good reason to, so I double-down on Bazel's only selling point while ignoring everything".
It's ok if anyone just wants to try stuff without having any rational or coherent argument to justify it, but let's not call that "problems".
Nonsense. All build systems implement a DAG of build targets.
> why caching tools can't be used effectively
Nonsense. There are a myriad of caching tools freely available, from ccache, sccache, build cache, etc. To get them to work, all you need to do is pass compiler launchers to them, which means setting a single env/build flag.
The author is either trying hard to make a mountain out of irrelevant molehills, or is outright being disingenuous.
Bazel is being sold as a solution desperately grasping at straw-like problems, while being feature incomplete.
And the example they gave is one that doesn't work so well.
> ccache, sccache, build cache, etc
The example they gave included the build cache misfiring. ccache doesn't work with Rust. sccache does, but can't be used with incremental compilation or when building binaries, making it essentially useless for what they needed, i.e. speeding up CI and local dev.
I really have no idea why you're still avoiding reading the article you're dismissing.
Yeah, that's an effective way to sell someone's promotion, but it's always a promise that is never met and instead it's a constant source of problems.
Sometimes reinventing the wheel just buys you a wobbly, squeaky wheel.
In our monorepo, we ended up just using a Makefile, since it basically does that. Of course the Makefile is now a few hundred lines and we have about 60 *.compose.yml files, but that's another story.
Our current solution still uses the Makefile, but we basically split the monorepo into a Python part and a TypeScript part, and each of those has its own monorepo build system (Yarn workspaces for TS, and Poetry with Pants for Python).
Other big complaints were:
- big CPU hog
- bad IDE integration (esp scala/java for intelliJ)
- it lets people write i-cant-believe-its-not-python, which is far too much rope to hang yourself with. I've seen horrors.
"Cargo is not a build system" reminded me of "Cabal is not a package manager"
https://ivanmiljenovic.wordpress.com/2010/03/15/repeat-after...
I'm using `crane` and `buildRustPackage` to build wasm front-end that is embedded into server binary (also rust). It works like a charm.
It seems like the ideal way to mix nix and cargo is to simply trust and delegate to cargo once you reach a rust dependency, not try to nixify it.
I think there was a rustup library for nix that probably would allow doing that but I haven't spent enough time trying (it's such a pain to debug flake.nix files)
You definitely should be leaving a sort of boundary between them. You can use fenix in your devshell to get a toolchain and then use crane for building. You just use cargo like normal when developing, including dependent crates, but CI can be handled with nix from that point.
I've got my flakes organized in a way that they are basically modular, so I can easily reuse files from other ones. If I need a rust toolchain, cargo builds, python toolchail, postgres, or whatever I can just import the file and change it up slightly (mostly names)
> This project is not yet polished. We are continuing to develop it in the open, but don't expect it to be suitable for most people…
So it’s not ready yet to be used for anything serious I assume.
If you want stability, Bazel is still a much better choice with their LTS release program.
As with all things Bazel, it's great when it works, but it has a steep learning curve that could be hard to justify in many situations. You need at least one person on the team who is willing to really spend time on it.
I successfully migrated recently and it's such a gigantic improvement.
Edge cases in Bazel are about as documented as any random fish you might find in a crack at the bottom of the ocean. And the most hilarious thing to deal with is "yeah this was written for a Google project, that other use case didn't apply to our project".
When eventually after the Xth issue I was told "why don't you just fork this standard package" (it was for typescript support inside the Rust repo) I was like "f*** bazel" and never touched it again.
It's amazing when it works. If it doesn't you're f******.
Some ex-Googler may want to use Bazel and may have very good reasons for it, but they’re often unprepared to take over the support duties that, at Google, are provided by the developer tool teams.
Could you explain how it beats for example, a Dockerfile in terms of what you need to write/get your hands dirty with (syntax, tooling, etc.) if your output goal is just an OCI image format to publish/run in container orchestration software? Is that too simple of a usecase?
Not only is that a very narrow set of use cases (Read: they only deploy statically linked binaries on mostly unix-y environments), Even the fixes for those very basic use cases take for ever.
Eg. This 7 year old issue is still open: https://github.com/bazelbuild/bazel/issues/1920 . To be able to create a statically linked library, we had to use: https://github.com/hotg-ai/librunecoral/blob/master/runecora... . Had to use some weird hack to build shared libraries too. Overall, it was just annoying.
TBH it’s hard to say which operate model is better anyway. They are simply incompatible :)
Related topic: https://youtu.be/Pzl1B7nB9Kc
https://github.com/bazelbuild/bazel/issues/1920
https://stackoverflow.com/questions/32845940/symbols-from-st...
We just had the ugly requirement of having to build a rust library/crate that uses C++ libraries (tensorflow lite + libedgetpu). That forced us to use bazel, on/for platforms that i have little experience / love for (Windows, Apple platforms).
Hitting the platform annoyances of each of those platforms just added to the pain. (On Windows, the 260 character path limit + on iOS Apple messing with ar binary used to create static libraries).
Left a very bad taste of trying to use bazel as a build tool. Even handwritten 1000+ lines of makefiles/experimental build systems like qbs did not cause so many headaches.
to show how awesome bazel is - in one step it can build (compile) and package them into a .zip file ready to deploy to NuGet-like repo or other place.
The fallacy here is that just because you are not at the scale of the FAANG does not mean that their tools/libraries/products are not suitable for you - e.g. it's not the case that only if you reach their scale, then maybe you should switch there. This is bs.
Similarly, they might choose some other tool out of google because it doesnt extract such a high price
Exceedingly tedious to use and has poor support in Clion.
It is likely that I don't have the constraints that this company has. The last time I tried this in Java I had to replicate my package dependency graph in Bazel. It was exceedingly pointless.
I found an example repo for rust and Bazel and it is just as tedious as I recall https://github.com/laramiel/bazel-example-rust
0 - https://github.com/bazel-contrib/rules_jvm/tree/main/java/ga...
a BUILD file declares a "package" and the demarcation line happens when it sees another BUILD file in a sub-directory. Then that's another "package".
Where it gets harder is when this language feature is integrated with an existing build system. C# is normally build with it's own build system, and to build it with Bazel you have to essentially rewrite some of it's logic in Bazel, which can be hard, but doable.
So you need several layers of tooling just to get something built.
It's bad
But really, I think that people don't go extremely granular on this. For Rust projects that are spread over 10-100 crates, I think that most people would just model the crate dependencies rather than at a file level. Mostly for pragmatic reasons, but also because that will likely model your test suite closely and your build requirements.
Bazel is a build tool, so one "module" (in Bazel you talk about target) per build artifact basically aligns with what people want. But you can in theory go really granular, making it so that tests only run if files it depend on change (less granular and that becomes "crates they depend on change")
> For Rust projects that are spread over 10-100 crates
It's already 10-100 crates you need to manually declare and keep track of, or use a yet another tool to generate this for you.
And for big projects you'd also want to go more granular and split by parts of your app I guess.
The "obvious" thing would be to either generate cargo files from Bazel stuff, generate Bazel stuff from cargo files, or have a third source that generates both. And like... yeah, if you are suffering from compile time issues, Bazel can make that stuff order of magnitudes faster (especially if your dep tree is pretty horizontal), so there's a reason to do it!
But fear not, (and I'm not sure if BAZEL does this yet, but does it for java). If you've declared only "a.cpp", but forgot "a.h" - then bazel would detect that (/showIncludes in MSVC, or file system sandbox in linux, etc.) - and it'll tell you - you need to add "a.h" to your BUILD file.
Fear not again, it's possible to create a BUILDIFIER command that given what BAZEL told you it can do it automatically for you. Hence the BUILD files needs to be kept simple, for everyone to understand (even non-engineers), while the ugly parts go into the .bzl files and workspaces.
So in tha java world (but any language could do it), anything that was discovered dynamically but not declared in, bazel can offer how to fix it (copy+paste the command offered, or tool makes it for you) - at worst report it and you have to fix it manually
The important thing is - bazel can use only the information provided in the BUILD files to detect - oh hey, things have changed in the CI - so I need to rebuild. Without relying on "cl/gcc/clang" to run this through preproc step to discover actual headers. or keep horrible make-like deps.
> This nix-alienation bifurcated our build environment: the CI servers built the code with nix-build, and developers built the code by entering the nix-shell and invoking cargo.
I don't anything wrong with this. Why would you throw the ability to simply use the Rust toolchain and all other tools directly?
It might be the fault of `nix-shell` which was a confusing and worked in some contrived ways. Nix Flakes have `nix develop` with an explicit "nix dev shell" construct which starts a shell that simply puts all the tools neccessary to build the project in the PATH (and possibly other relevant envs).
In our project developers simply need to install Nix and run `nix develop` (or use direnv) and from there on they mostly don't need to remember about Nix during their local work. Only when they need to touch something build/CI related they need to touch nix files, which is maybe 5% of the work, and then they can just ask for help.
The higher level tests do require a bit of thought so they reuse Nix building system, instead of fighting against it.
> Our security team chose Ubuntu as the deployment target and insisted that production binaries link against the regularly updated system libraries (libc, libc++, openssl, etc.) the deployment platform provides.
> Furthermore, the infrastructure team got a few new members unfamiliar with nix and decided to switch to a more familiar technology, Docker containers. The team implemented a new build system that runs cargo builds inside a docker container with the versions of dynamic libraries identical to those in the production environment.
I think at this point it was game over for using Nix.
Nix has a great support for building OCI (aka Docker) containers. It's great to be able to write 10 lines of Nix and just import all the required derivations from existing Nix build system and get a quick to build, nicely cacheable container that has only 35M.
But if you're not going to use it, then you turn a benefit into an extra work.
Using Nix to build inside Dockerfile is a terrible combination, eliminating lots of benefits of each done properly in separation.
> The idea of migrating the build system came from a few engineers (read Xooglers) who were tired of fighting with long build times and poor tooling.
It seems to me that it comes down to this. People didn't want to learn Nix, so they constructed bunch of incompatible ways of doing things.
Eventually there were already bunch of people familiar with Bazel so that's what they used. Had you have few engineers familiar with Nix, they would restructure things and make them work with Nix.
Some final tactical advice for people that would want to solve these problems with Rust. Use Nix Flakes, they are soo much better than old style Nix. Use `crane` for Rust - IMO the best compromise between "not rebuilding everything" and "simple wrapper over cargo". Use cachix or some shared Nix caching. Let developers work in `nix develop` shell and do their thing without thinking too much about Nix. Whenever they need part of codebase running reproducibly (fixtures for tests etc.) tactically invoke `nix build` and `nix run` . For normal Linux distros, provide OCI containers built natively with Nix. If you need raw binaries in order of preference: cross-compile to Musl (can't be done for C++), patchelf (a bit risky with C++ as OP mentioned), maintain a parallel build procedure via Dockerfile, but don't use Nix in it, and use it only to build the artifacts and nothing else.