When to use Bazel?
earthly.dev
earthly.dev
So the answer IMHO to “when to use Bazel” is “never” :)
Also, Google tooling knowledge doesn't transfer very well to outside projects. Everything is special on the inside there. They use an internal version of Bazel called Blaze that's totally integrated with everything, with entire teams dedicated to the tooling, and only a few officially supported programming languages, so of course it works smoothly.
I understood some of the design decisions and tradeoffs for why it works the way it does, but I didn't like it. I certainly wouldn't reach for it for a personal project. Possibly I still just didn't "get it". I've used dozens* of build systems at this point, but hey.
* possibly exaggerating
A good number of their open source projects fall apart in real-world use. If you visit their repositories you'll often get the sense that beyond working on interesting problems, their engineers have no desire to properly maintain these projects for the next few years. The gamble of relying on a Google product unfortunately also applies to their open source projects.
What about Golang? Angular? Dart? Even the bigger stuff like Chrome and Android while a bear to deal with are integrated with every day by a lot of companies.
What is some Google OSS that has failed in the way you mention?
Sure it wasn't great, but the few projects I did with Angular 6,7,8 was kinda nice. I still have an AngularJS project in production which is also nice :)
Of course any new JS project is now done in Svelte (nothing simpler) or as nice
PS. It's nice to be nice :D
and there still no guarantee of no such regressions one month later.
I assume it's OK if it's your only platform—though I got my programming start on Windows, with open-source languages, and distinctly recall how managing the same tools and builds got way easier and more reliable when I switched to Linux. Luckily WSL2 is getting semi-decent and can call out to Windows tools, so it's getting easier to standardize on one or another unixy scripting language for glue, even if some of your tools on Windows are native.
Part of the trouble with it, though, is that there are multiple ways to end up with some kind of linux-ish environment on it, and that all of them introduce quirks or oddities. You end up having to wrestle with stupid questions like "which copy of OpenSSH is this tool using?" It's probably less-bad if you're in a position to strictly dictate what Windows dev machines have installed on them, I suppose. Git is installed, but only with options X, Y, and Z checked, no other configurations allowed; WSL2 is installed and has packages A, B, and C installed; and so on, crucially preventing the installation of other things that might try to pile on more weirdness and unpredictability.
Those kinds of messes are possible on macOS or Linux or FreeBSD or what have you, but generally don't happen. I think it happens on Windows because every tool's trying to vendor in various Linux-compat dependencies rather than force a particular system configuration on users, or having to try to deal with whatever maybe-not-actually-compatible similar tools the user has installed. So they vendor them in to reduce friction and the rate of spurious bug reports. Basically, the authors of these tools are seeing the same problems as anyone trying to configure cross-platform development with Windows in the mix, and are throwing up their hands and vendoring in their deps, which compounds the problem for anyone trying to coordinate several such tools because there's a tendency for everything to end up with weird configurations that don't play well together.
Unpopular opinion but I think the Unix approach of lots of little tools with arcane configuration files all blaming each other sucks. I get why it exists, why it became popular, and why it remains popular in some scenarios, but I don’t feel beholden to it.
If we are talking about stuff like backend web development then I'd say no, it's actually a pain to work with - few people bother deploying to Windows these days, many libraries/frameworks treat it like a second class citizen, tooling doesn't translate well (or at all).
Even frontend webdev is very unixy due to node/npm and popularity of Macs in that space.
When you're talking about OpenCV, GPU related things, game dev then the situation is reversed.
So it highly depends on what you're developing.
Obviously this isn't a deal breaker in most situations, but it is a bit annoying. I've observed similar performance problems running Webpack on Windows in the past.
Also take a look at this talk on NTFS, file system speed, and how the presenter optimized the Rust installer: https://youtu.be/qbKGw8MQ0i8 . There is an awful lot that can be done to tools that deal with lots of files to perform well under Windows too.
No. It was the VFS design. They have lots hooks and extension points for 3rd party, there is no way to improve without breaking backward compatibility. It was the reason why many antivirus was broken in the first few releases of windows 10.
The only time I see Windows frontend devs is full-stack .NET shops or low cost outsourcing centers. If nothing else then just for practicality - you need to test on Safari and iOS so might as well get Mac devices.
OTOH, in practice, if you're using a lot of open source tools and have developers working on every single halfway-plausible OS except windows, things are probably pretty OK, but then you throw Windows in the mix and suddenly the time you spend supporting your builds & tools shoots way up. That makes it feel like it's Windows' fault.
$ file -L /usr/lib/git-core/git-rebase
/usr/lib/git-core/git-rebase: ELF 64-bit LSB pie executable, x86-64, version 1 (SYSV), dynamically linked, interpreter /lib64/ld-linux-x86-64.so.2, BuildID[sha1]=95d963a593ffeeb306e8525aacce02f1cbc7dabd, for GNU/Linux 3.2.0, strippedIt was a complete waste of Mercurial developer's time that would have been better spent creating a GitHub equivalent.
If it doesn't come from Microsoft, Windows developers don't care.
Learn that lesson or repeat the mistake.
In my experience, most software development tools don't work on it, at least not as well as on Unix-like systems. But I'm not a game dev.
There's a standard find_package/find_library macro and definitions for OpenCV (which is your use case here). Then use either Microsoft's C++ package manger or Conan to actually install the libraries.
IMHO, Bazel is a classic example of Conway's law[0], and it falls squarely in the "big corp" class of software. You have to be running into issues like 6+ digit CI compute costs or latency SLOs in hundreds-of-teams monorepos before Bazel really starts to make sense as a potential technical solution.
But where it does work well, is when you have a large complex codebase in a monorepo, and you need reliable caching to keep build times down.
> I just `find . -name "*.h"`
That sounds like it would cause problems with not triggering a build if a .h file is deleted. And would trigger unnecessary builds if any .h file is modified. But if your builds are that fast, maybe that isn't a problem.
ccache is going to be as fast as any checksum dependency tracker --- otherwise you're relying on a timestamp, which isn't any better than "bare" Make. We use `*.d` rules from the `-MD` flag to prevent the "deleted header" case.
Considering the brevity of the Makefile (it looks like a "basic" build), performance, and the practical robustness, it's in a pretty great trade-off space.
2. Make lets you define inputs and outputs, but doesn't enforce them. Bazel sandboxes build actions; if you try to import a dependency that you haven't listed as a direct or transitive dependency, the build will fail. This is leads to users never (or at least very rarely) having to run `bazel clean`.
To take it a step further. I never left the "comfort" of bash-build and bash-deploy scripts.
In most of my projects (pro and personal) there is a deploy-aws-prod.sh and deploy-aws-dev.sh (or some variation)
None is longer than a 10-20 bash commands.
These "projects" are webservice, ml-models, batch-processing(think distributed clusters)
It's ugly and perfect-enough at the same time.
YMMV
This is becoming more and more of a pet peeve of mine, about to become an outright annoyance; the adding or changing of tools without things actually improving, or the purported improvement not actually being relevant or even close to the most important thing that time should be spent on.
https://mcfunley.com/choose-boring-technology is a good starting point for moving away from this mindset.
I had a Go project up until earlier this year, I could set it up myself. I picked makefiles to build it (not bazel, that would be overkill; not mage, that would mean more coding), and it worked just fine. I never had a compelling reason to move away from it.
Actually it was more than fine, it was a relief; my previous experience with tools like that had been Maven (where everything is XML and you need a plugin to do basic things like move or remove a file) and the Javascript ecosystem, from before everything was merged into package.json / nodejs-style. Dependency management too.
Boring technologies are easier for people better and making due with what they have and accepting sometimes very large limitations.
ADHD makes boring tasks harder for instance, and boring technologies give you no escape from the tedium.
With a little fancier technology though, you at least have hope ;)
I will never use bazel again. Not worth the effort required, and doesn’t work out of the box as advertised.
Get back to me when you work for a company with a monorepo that has to build and test everything in CI for every change (taking something like 200 CPU hours) because they didn't have the foresight to use Bazel.
In practical terms, the makefile or whatever is usually good enough.
Bazel isn't designed to make things easy. It isn't even designed to make things fast (though it is easy to add that on due to repeatable builds). Bazel is designed to prevent you from shooting yourself in the foot; providing strong controls on project dependencies, both dependencies of your project as well as other projects' dependence on your project.
That means in 99% of companies the answer is no. You're better off with CMake or something thing like that.
That's a very rare occurrence in software. Honestly it's probably about as common as in housing. You can rebuild parts of a piece of software but the build system is generally really foundational and almost never changes in my experience. Especially if the software isn't a library so nobody except the main developers has to interact with the build system.
What's often impossible to replace without a full rewrite is a backend's DBMS, and people tend to underestimate how tight of a coupling that is. Only gets worse with abstraction.
That smells like covering up an architecture problem (tight coupling) with a build tool. Not a recipe for success.
CI tooling is orthogonal; your CI tooling can invoke Bazel or Makefile+custom caching, it wouldn't really matter. But the problem in CI is that you lack state; whereas a developer's machine might e.g., have a bunch of .o files from a previous build, CI will be starting from a clean slate¹. So builds take longer. My understanding of Bazel is that it has decent support for caching and can thus makes those rebuilds a lot faster, which it has b/c it understands exactly the inputs to the steps at hand. (Which, if you write your own caching layer, you'd have to figure out.)
… but … I've also seen the same thing the top commenter has, with Bazel: endless suggestions to use Bazel … but migrating a large existing codebase is pain, and if you're not willing to put your work where your mouth is … then the status quo prevails.
¹While here CI lacking state is a sort of negative b/c of the time we need to rebuild it, normally, I think this is a good thing: it means you're constantly verifying you can build from scratch. Or, at least, something close to scratch, modulo things like vendoring. (I.e., if you do an "apt-get" from Ubuntu's servers in your CI pipeline … well … your build obviously isn't quite from scratch.)
It makes more sense when you realize it was really built to deal with statically linked C++ - there's no guarantee of ABI, so you really need to ensure that all your dependencies are built the same way as your project to ensure compatibility (and the dependency graphs in monorepos get DEEP).
For any language with a stable ABI that utilizes dynamic linking, the complexity is far more then the benefit. The whole monorepo concept and all the tooling that goes with it really makes more sense when you think of it in that context of statically linked code that can't easily be packaged into something as portable as a JAR, so it's critical to build it all the same way.
That's a perfectly fine choice if you have bodies to throw at the problem and can deal with the consequences of that choice, but most people can't. Google can and so it does. If it wasn't as successful as it is, it probably would have collapsed under the strain of the people and systems required to manage that code.
Other tooling decisions place a burden on every SWE. Their config language is a weird trash that everyone has to sorta learn. Their deployment and integration testing tooling have similar qualities, and it's not cause of the monorepo. One consequence is that many teams cannot feasibly deploy microservices, even though it'd often make sense to, which is even worse if they're using C++.
I’m really glade that Meson and Ninja exist however. I hate GNU Make like few things in my toolbox. M4 really is an awful language.
Specifically, Bazel is really at home with Java/C++ and Linux. Sure, it kinda works elsewhere, but you should be considering other options.
We very happily runny a polyglot monorepo w/ 5+ languages, multiple architectures, with fully reproducible artifacts and deployment manifests, all deployed to almost all AWS regions on every build. We update tens of thousands of resources in every environment for every build. The fact that Bazel is creating reproducible artifacts allow us to manage this seamlessly and reliably. Clean builds take an hour+, but our GH self-hosted runners often complete commit to green build for our devs in less than a minute.
The core concept of Bazel is very simple: explicitly declare the input you pass to a rule/tool and explicitly declare the output it creates. If that can click, you're half way there.
Yes, and it's very important to note that Bazel does nothing to solve the problem about having a reproducible and hermetic runtime. Even if you think you aren't linking against anything dynamically, you are probably linking against several system libraries which must be present in the same versions to get reproducible and hermetic behavior.
This is solvable with Docker or exceptionally arcane Linux hackery, but it's completely missing from the Bazel messaging and it often leaves people thinking it provides more than it really does.
In particular, it’s almost impossible to get glibc to link statically. You need to switch to musl or another libc that supports it.
Go creates static binaries. I was a bit bummed out to learn that Rust doesn't by default since it requires glibc (and copying a binary to an older distribution fails because glibc is too old).
Two commands to get you running:
1) rustup target add x86_64-unknown-linux-musl
2) cargo build --target=x86_64-unknown-linux-musl
One thing Go has going for it is that they have a more complete standard library so you can do without the pain of cross-compiling—which is what you're doing if most of your libs assume you're in a glibc environment but you really want musl.
I know because I recently tried compiling a Rust application that had to run on an air-gapped older Debian machine with an old glibc, and the easiest solution was to set up a debian container, download Rust and compile there, instead of fixing all the cargo deps that didn't like to be built with musl.
So I just migrated everything over to that (which consisted of enabling the `rusttls` feature for every crate) and made another build with musl. Everything worked perfectly fine, and since it's not a performance sensitive application, there was basically no drawbacks. The binary became a bit bigger, but still under 4MB (with other assets baked into it) so wasn't a big issue.
For other languages I'd say openssl is probably the next most common.
I don't think that's true for Rust. It defaults to depending on glibc, at least on Linux, and you need to explicitly build with musl to get static binaries.
There is no platform ABI mandate nor Linux kernel requirement for userspace applications to be dynamically linked to anything. If you can talk to the kernel, you can do anything those libraries can do.
On some architectures, atomics are implemented in the kernel by the vDSO shared library (__kernel_cmpxchg), which is supplied to user via libatomic. You can ignore it, but then you cannot interoperate other code (any which uses libatomic.so) without introducing data-races into the code which were not present in the shared-library version, since they may attempt to execute their atomics differently and thus incorrectly.
Maybe I just like cults ^_^
Also, I do prefer to avoid dynamic linking altogether.
I dread running builds because it takes like ten minutes to run `make all`. I spent like four hours writing Bazel build files for it and all of a sudden my clean rebuilds were taking like five minutes and my incremental builds were taking a couple of seconds. It was fantastic.
Ultimately it was rejected because other people on the project hated working with bazel and weren't familiar enough when it to debug any problems. If a build didn't work, the only solution was to call me because no one else was used to interpreting the bazel errors: if a curl command failed in a script in a docker file somewhere, everyone knew what that meant. But I was the only one who would see "IOException fetching... Error 401" and immediately know that they provided the wrong password.
I also consulted for another project where people were using containers as VMs, running systemd and copying in updated binaries, because they, too dreaded rebuilding the containers. Bazel, again, made it so that they could rebuild everything in a matter of seconds. That project is still happily using bazel last I checked.
> B̶a̶z̶e̶l̶Docker uses whatever toolchain is laying around on the h̶o̶s̶t̶Internet
More precisely, Docker can use exact, well-specified, cryptographically-tamper-proof environments (images); but the only way to actually specify and build such an environment is via shell scripts (in a "Dockerfile"). In practice, such scripts tend to run some other tool to do the actual specification/building, e.g. `make`, `mvn`, `cargo`, `nix`, etc.
If those tools aren't reproducible, then wrapping them in Docker doesn't make them reproducible (sure we can make a snapshot, but that's basically just a dirty cache). In reality, most Dockerfiles seem to run wildly unreproducible commands, e.g. I've encountered Dockerfiles with stuff like `yum install -y pip && pip install foo`.
If those tools are reproducible, then wrapping them in Docker is unnecessary.
Also note that outputting a container is nothing special; most of those tools can do it (e.g. in a pinch you can have `make` run `tar` and `sha256sum` "manually")
For this reason we are building a Bazel competitor with a less restrictive build environment which can also be applied to projects with lower than 1M line of code.
For retro game projects, the core of your game might be written in C or C++, which you want cross-compiled. That's easily within reach of stuff like makefiles. But then I start adding a bunch of custom tooling--I want to write tools that process sprites, audio, or 3D models. These days I tend to write those tools in Go. I'm also working with other people, and I want to be able to cross-compile these tools for Windows, even though I'm developing on Linux or macOS.
My Bazel repository will download the ARM toolchain automatically and do everything it needs to build the target. I don't really need to set up my development environment at all--I just need to install Bazel and a C compiler, and Bazel will handle the rest. I don't need to install anything else on my system, and the C compiler is only needed because I'm using protocol buffers (Bazel downloads and compiles protoc automatically).
Once you have bazel, you can distribute the workload so that you could have thousands of machines each producing artifacts to be shared with other build machines.
Then you can set up your dev machine to rely on those caches, so your local builds either use everything directly from cache or instruct a remote builder to produce the artifact for you.
No matter what you change, because the dependencies are graphed precisely, you only need to rebuild a very tiny set of artifacts impacted by your change.
Of course, this doesn't actually work in practice.
Your builds probably aren't deterministic, so your graph of artifacts won't be either, causing lots of stuff to rebuild. Also, it may work great for something like Java that produces class files, but not provide any caching at all for ruby.
Debugging and supporting it is a full time job for a team of engineers that dramatically outweighs the cost of keeping your projects sensibly sized.
You might think that as the project gets larger, it's totally worth it to use bazel for that sweet caching. But in reality the graph construction and querying will become so bloated that just figuring out which targets need to rebuilt becomes a full time engineering effort that breaks constantly with all tooling upgrades.
Also, the plugin ecosystem is just poor.
Bazel is the perfect storm of computer scientists loving big graphs, Google exporting an open source project and then rebuilding it internally, and inexperienced engineers being sold on a tech as being obviously right because all the big players use it.
But if you actually have to get something done for your business to exist, it's a lot of unrelated work to keep the beast fed and happy.
I guess it's theoretically possible, but why would you do this? This sounds like some kind of "I want to prove it can be done" project like running a web server on a Commodore 64, or making a kitchen knife out of cardboard. Yes, you could figure out a way to download and install dependencies using Make and Curl, and maybe you could build the same library for multiple targets using a bunch of variable expansions, possibly by invoking make multiple times, or including the same makefile multiple times. Make sucks; it sucks a lot; I'd be miserable.
And then if you aren't very careful, you'll end up forgetting to declare some dependency, or some flag change will invalidate your build but Make won't do anything, or you'll make a change to your makefile and forget to clean, and then you spend another hour or day debugging some build that was made out of stale parts. I've done this before, which is why I avoid make for everything but the smallest projects.
> without java dependency
Bazel does not have a Java dependency. I think you might be assuming that because Bazel is written in Java, somehow that means that Java must be installed on your computer. This is not true. Bazel is a self-contained executable you can drop in /usr/local/bin.
I really don't understand your perspective here at all. I've seen people defend make, but make is full of so many traps and gotchas--I wonder how someone could use make and then decide that these traps and gotchas are okay. The projects that successfully use it tend to be small projects, use generated makefiles (automake, cmake, etc), or tend to be a bit simpler.
I use make myself, but as a rule of thumb, only for projects with a few files.
> what real extra benefit Bazel brings in here
Builds are hermetic by default, so unless the developer chooses to escape the sandbox, everything is guaranteed to build on other machines with no additional setup.
(Also, I genuinely hate when I have to manually install build dependencies system-wide and pray that there will not be any conflicts. Having everything pinned to specific sha256 or git hashes by design is a breath of fresh air)
The parenthesized comment is funnier to me because I had to download a specific bazel version to build.
I'd expect Tensorflow to have some non-hermetic build actions, but if choosing a specific Bazel version was the only thing that was required to build it, that's awesome!
If you use bazelisk to provide your `bazel` command, it'll download the appropriate Bazel version for the repo you're trying to build.
So, spurious changes (touching a file) will result in a cache hit, while hidden changes (changing an environment flag used by Make) are caught.
This is particularly important if verifiable builds are needed for SoX compliance.
https://github.com/bazelbuild/bazel/blob/34ce6a23f5a2be58bb5...
I don't know enough about bazel to answer your question definitively, but "stamping" is what you want to search for.
It checksums the preprocessed files, so neither `__DATE__`, no comment changes should affect your build times.
The .a/.so/.exe would no longer match the inputs (the .o files), causing a re-link.
Bazel can also read source file hashes from the filesystem if the filesystem supports it.
That means you can't have missing dependency links (very common and hard to debug in Make or CMake for example). You can know what outputs a change affects so you can only build/test a subset of the project in CI. Incremental builds become totally reliable. Etc.
Any makefile doing anything remotely complex is definitely not readable later.
https://github.com/depp/bazel-gba-example
Cross-compiling toolchains are a pain to set up in Bazel. Once you get them working, it's nice.
This somewhat long article hopes to answer questions about Bazel for future people in a similar position to me. If you know a lot of Bazel then you might not learn much but if you’ve vaguely heard of it and are not sure when its the tool that should be reached for I’m hoping this will help.
I'm also surprised at how little experience everyone had with bazel alternatives. I wonder if buck or pants is easier to work with.
Oscar in the article had experience with Pants but preferred Bazel and Bazel certainly has the most usage at this scale.
What documentation exists is flawed and what isn’t flawed has massive holes. It was so bad that I had to ask a few Google employees if there was secret internal documentation somewhere. There isn’t, and, they almost all hated using it as well.
I went back to Make and had the whole repo building in an afternoon.
Bazel is a great idea but like most other Google OSS projects it doesn’t have strong enough documentation to form a community or enough of a community to create good docs.
Even if I brute forced my way into using Bazel, I couldn’t ask an employee to learn it.
I hear that Blaze is much easier to work with.
We use Bazel’s rules_docker as well, and I would caution someone evaluating it with a note from out experience.
What Bazel does well (and as well as Bazel fits your use-case) Bazel does extremely well and is a reproducible joy to use.
But if you stray off that path even a tiny bit, you’re often in for a surprisingly inexplicable, unavoidable, far-reaching pain.
For example, rules_docker is amazing at laying down files in a known base image. Everything is timestamped to the 1970 unix epoch, for reproducibility, but hey, it’s a bit-perfect reproduction.
Need to run one teensy executable to do even the smallest thing that’s trivial with a Dockerfile? Bazel mist create the image transfer it to a docker daemon, run the command, transfer it back... your 1 kb change just took 5 minutes and 36 gb of data being tarred, gzipped, and flung around (hopefully-the-local) network.
It may not be a dealbreaker, and you may not care, but be forewarned that these little surprises creep up fairly ofen from unexpected quarters!
Edit: after 2-ish years of Bazel, I would say that for 99% of developers and organizations, the most likely answer is "never".
It's so so much worse than the homegrown software it replaced.
Hermetic sounds nice in theory, but doesn't actually matter and comes with a massive cost.
1. Bazel is slow. Like really fucking slow compared to your native build tools. Startup time can be insane if you use multiple rules (languages) in your repo. There's a great feeling when you get to work in the morning to build your go project but you have to install the latest python, nodejs, and ruby toolchains because someone updated them on master. The cache works well, except with a 1000 devs something will always be invalidated on master.
2. The documentation sucks. It's written to explain concepts with how things work, with no examples of how to actually complete tasks you care about. Of course that wouldn't help either because the bazel setup you're using is heavily customised.
3. Everyone on your team now has to learn yet another DSL, and a bunch of commands to run. I would easily spend 5 hours per week either waiting for, or debugging some issue in bazel.
All this for what benefit? Everyone's on the same version of a dependency? Not even sure this is a desirable property.
Also the monorepo is slow as jelly to work with and many tools or editors struggle to open it. Good luck getting code completion to work well too.
It's one of those ideas that are nice in theory, but awful in practice. It's possible that with enough manpower you may be able to use it effectively, but we certainly were not.
Probably because its written in Java.
What sane person writes Java in the 2020s ?
I mean, aside from having to dance the JVM maintenance dance, you're also faced with the fact that Java is a memory hog unless you spend half your life tweaking obscure parameters in XML files.
The sooner Java dies the better, frankly. Java was great back in the day, even ahead of its time, but today it's just a relic.
Use Go, use Rust ... hell, practically anything is better than Java.
Would it benefit from being compiled with GraalVM?
Java itself is a language, and cannot be slow or fast. However, the most popular (and thus well-supported) implementations of Java are slow. On those implementations, the idiomatic way of writing Java leads to massive amounts of memory allocations, which force Garbage Collector to allocate much more memory than it needs and perform complex collection strategies. Many allocations also destroy memory locality and completely negate the effect of CPU cache.
> Would it benefit from being compiled with GraalVM?
Maybe, maybe not. But GraalVM, at least for now, is not the standard way of running Java programs.
I think the point of the parent comment was that it's bad to choose Java for anything other than long-running server-side applications (where JIT compilation with Hotspot might actually prove valuable for squeezing out those last few percent of throughput). For anything else, Java is really a sub-optimal choice - the startup time is horrible, the speed is not really great, and the JVM maintenance is painful. For a build tool, which is invoked many times to do relatively small amounts of work (compared to long-running server-side applications), it just doesn't make sense.
Bazel makes use of a client-server architecture. Java is used for the server component, for which startup time is not an issue. The CLI client is written in C++.
In that case my last paragraph should be discarded.
Bazel takes on dependency management, which is probably an improvement for a C++ codebase where there is no de-facto package manager. For modern languages like golang where a package manager is widely adopted by the community it's usually just a pain. e.g Bazel's offering for golang relies on generating "Bazel configurations" for the repositories to fetch, this alternative definition of dependencies is not what all the existing go tooling are expecting, and so to get the dev tooling working properly you end up generating one configuration from the other having 2 sources of truth, and pains when there's somehow a mismatch.
Bazel hermeticity is very nice in theory, in practice many of the existing toolchains used by companies that are using Bazel are non-hermetic, resulting in many companies stuck in the process of "migration to Bazel remote execution" forever.
Blaze works well in Google's monorepo where all the dependencies are checked in (vendored), the WORKSPACE file was an afterthought when it was opensourced, and the whole process of fetching remote dependencies in practice becomes a pain for big monorepos (I just want to build this small golang utility, `bazel build //simple:simple` and you end up waiting for a whole bunch of python dependencies you don't need to be downloaded).
And this is all before talking about Javascript, if your JS codebase wasn't originally designed the way Bazel expects it you're probably up for some fun.
Conversely, as many comments here observe, it is terrible as a small build system. There, you want to be able to get started quickly and pull in dependencies without thinking too hard. A simple but easy to understand approach (even one based on make) might work. This is just a different problem than what Bazel solves.
My personal opinion (and I should emphasize, I am absolutely not speaking for anyone here, least of all Google) is that the best way forward is to take the ideas of Bazel (hermetic and deterministic builds) and package them as a good small build system, perhaps even compatible with Bazel so you don't have to rewrite build rules all the time. I also think compilers and tools can and should evolve to become good citizens in such a world. But I have no idea how things will go, it's equally plausible the entire space of build systems will continue to suck as they have for decades.
[1]: http://neilmitchell.blogspot.com/2021/09/small-project-build...
[2]: https://neilmitchell.blogspot.com/2021/09/reflecting-on-shak...
[3]: https://neilmitchell.blogspot.com/2021/09/huge-project-build...
It's on my list of things that I will inevitably never get to.
How does https://please.build measure up?
It's great at being able to offload work to a remote server, but in my opinion that should never be the only way you can get work done. Local development should always be the default, with remote execution being an _option_ when available.
As far as criticisms go, “we lost a few days of work when a global pandemic shut the world down” seems pretty minor.
That said, I too am a fan of offline development and I find that bazel makes that pretty easy for me. I just run bazel fetch before I get on an airplane.
The existing codebase I was working with did not lay out its dependencies and files in the manner expected by Bazel so dealing with dependency/include hell was frustrating.
Then there was a large portion of the project that depended on generated code for its interfaces (ie. similar to protobuf but slightly different) and trying to dive into the Bazel rules and toolchain with no other Bazel experience was not fun. I attempted to build off of the protobuf implementation but kept finding additional layers to the onion that didn't exactly translate to the non protobuf tooling. The project documentation seemed out of date in this area (ie. major differences in the rules engine between 1.0 and subsequent versions) and I couldn't find many examples to look at other than overly simplified toy examples.
All in all a frustrating experience. I could not even get far enough along to compile the interfaces for the project.
I think part of the issue was that the generated code from the 3rd party tool did not align with the Bazel expectations about code layout (ie. paths relative to WORKSPACE). Likely solvable but not in the amount of time I had available to dedicate to the effort.
The good:
- Amazing for Go backends. I can push a reproducible Docker image to Kubernetes with a single Bazel command. I can run our entire backend with a single command that will work on any developer's box.
- Amazing for testing. All of our backends tests use a fully independent Postgres database managed by Bazel. It's really nice not having to worry about shared state in databases across tests.
- We can skip Docker on macOS for development which provides on the order of a 10x speedup for tests.
- BuildBuddy provides a really nice CI experience with remote execution. Bazel tests are structured so I can see exactly which tests failed without trawling through thousands of lines of log output. I've heard good things about EngFlow but BuildBuddy was free to start.
- Really nice for schema driven codegen like protobufs.
The bad:
- Bazel is much too hard for TypeScript and JavaScript. We don't use Bazel for our frontend. New bundlers like Vite are much faster and have a developer experience that's hard to replicate with Bazel. Aspect.dev is doing some work on this front. One large hurdle is there's not automatic BUILD file dependency updater like Gazelle for Go.
- Windows support is too hard largely because most third party dependencies don't work well with Windows.
- Third party dependencies are still painful. There's ongoing work with bzlmod but my impression is that it won't be usable for a couple of years.
- Getting started was incredibly painful. However, the ongoing maintenance burden is a few hours per month.
I built this[0] for the Bazel + yarn setup that we use at Uber. We currently manage a 1000+ package monorepo with it.
Curious wouldn't `go run` give you the same? pure go code is supposed to be portable, unless you have cgo deps I guess?
> I can push a reproducible Docker image to Kubernetes with a single Bazel command.
That's definitely an upside over what would otherwise would probably default to a combination of Dockerfiles and scripts/Makefiles, does it worth bringing in the massive thing that is Bazel? depends I guess.
I'm curious: would you say your experience with golang IDEs / gopls is degraded? did you do anything special to make it good? I often feel like development is more clunky and often I just give up on the nice-to-haves of a language server e.g often some dependencies in the IDE aren't properly indexed, I can probably get Bazel to do some fetching, reindex and get it working, but it will take 3-4 minutes and I just often choose to live with the thing appearing as "broken" in the IDE and getting less IDE features.
We use protobufs and pggen. Bazel transparently manages the codegen from proto file to Go code.
> would you say your experience with golang IDEs / gopls is degraded?
Yes, that's our biggest pain point with Go and Bazel. I haven't been able to coax IntelliJ into debugging Bazel managed binaries. To enable IntelliJ code analysis, we copy the generated code into the src directory (with a Bazel rule auto-generated by Gazelle) but don't add it to Git.
I've tried the IntelliJ with Bazel plugin a few times but I've always reverted back to stock IntelliJ.
I would love to know more about this. We use bazel and Postgres, but there is no Postgres isolation between tests currently which is very predictably a large source of flakyness, and also a blocker for increasing test parallelism.
I’ve been scheming on perhaps using pg_tmp orchestrated outside of bazel, but if there’s a way to do everything in bazel that would make the transition way easier.
I posted this gist a while ago: https://gist.github.com/jvolkman/6e61c52d953677f66f32d8f77b1...
That said, this requires one postgres instance for each test, which is pretty heavy. So we keep our tests pretty coarse at the Bazel level (i.e. one Bazel test is actually 10+ Python tests). I've thought about other approaches, such as using a Bazel persistent worker to wrap test execution and manage a single Postgres instance with empty databases created from templates, but never got around to trying anything.
In addition to the techniques in the linked comment, I'm currently investigating managing the entire Postgres data dir as a Bazel output directory. That way, we avoid executing the schema for every test. Instead, we'd start a new Postgres cluster against the existing data dir without re-executing the schema or creating from a template DB.
Integration testing is a breeze through data dependencies. The reproducibility guarantees means we can reference container image SHAs in our Terraform and if the image didn't change the deploy is a no-op.
Bazel is an outstanding build system that handily solves a lot of practical problems in software engineering. Not just "at scale". Practical problems at any scale.
https://github.com/bazelbuild/remote-apis
In my opinion the protocol is fairly well designed.
https://sluongng.hashnode.dev/
Tools like Earthly (and other BuildKit HLB frontend languages) will help more teams get some of the main benefits of Bazel with a lower bar of complexity. If you already know how to write Dockerfiles and Makefiles, you can write Earthfiles without much additional learning necessary. It provides a build that is repeatable, incremental, self-contained, never dirty, shared-cacheable, et cetera.
See https://github.com/bazelbuild/rules_go/issues/1486
Not very confidence inspiring when Google’s build system falls over when you combine three technologies that are used commonly throughout Google’s code base (two of which were created by Google).
If you’re Google, sure, use Bazel. Otherwise, I wouldn’t recommend it. Google will cater to their needs and their needs only — putting the code out in the open means you get the privilege of sharing in their tech debt, and if something isn’t working, you can contribute your labor to them for free.
No thanks :)
Its very flexible, but in this case the flexibility is simple: you specify your input globs, output globs and task dependencies for each package.json command you want it to support and it will use that as the basis to detect cache invalidation.
Works rather well.
If possible, start fresh.
I don’t see myself working without it from now on.
Bazel is not just a build toolchain it is a test toolchain as well. The problems it solves with tests with respect to running only relevant & applicable tests to the changes being pushed is what many people have dreamed about before or came up with some crude solution to approximate the dependencies. If my test did not print out `(cached)` I know something in the dependency graph was disturbed in some way that may be completely not obvious to me, especially in a gigantic repo.
Bazel query is also worth mentioning to visualize dependencies with its xml output and which can also enforce architectural barriers asserting empty `somepath` queries between two targets that should not meet in the middle somewhere.
It would be neat one day if bazel integrates in some semantic logical level of dependency checking beyond the physical but that may be too expensive of an operation.
It's just a graph of tasks that take inputs in, and produces artifacts out. So long as the inputs to a task don't change then it can be cached.
The caching is indeed very good if you set it up properly.
In the hands of a level 5 gradle wizard it can also be extended to do ANYTHING, with the usual caveat that you may experience regrets when your gradle wizard leaves.
Source: We (MAANG) dumped the thing and rewrote it from scratch.
I've worked at Google for a long time. Every time I get a chance to work on something outside Google's codebase, it's like a breath of fresh air.
I wouldn't recommend Bazel or Blaze to anybody. Sorry to the people who wrote it.
Builds at Amazon were definitely not problem-free, but they were different problems that don't really resemble that description.
Then I started a project this year that will be a Rust, Swift, Kotlin, JS codebase (maybe C# too if Rust for Windows isn’t stable enough) and I had a look at Bazel again, but already the official tutorial from last year didn’t work on the latest stable release so I gave up on it again. Anyone had similar experience and pushed through? Is it worth it?
I currently have a bunch of Python scripts and will most likely end up with Gradle which is ok but not great in my experience.
What is supported needs to be inferred from this file, as far as I can tell: https://github.com/bazelbuild/rules_rust/blob/main/rust/plat...
I haven't used it myself so no commentary, I just have a friend who started there recently so it's on my radar
https://www.buildbuddy.io/ https://www.ycombinator.com/companies/buildbuddy
See this example: https://github.com/driftregion/bazel-c-rust-x86_linux-armv7_...
Our CI consists of tests, builds and deployments (all docker containers) that are kicked off in bitbucket pipelines. Is there some way that either our local development or CI builds can be improved by using earthly? Appreciate any info!
Edit: enjoyable to a point that i half considering doing contracts for bazel based build systems.