Build systems a la carte: theory and practice (2020)
simon.peytonjones.org
simon.peytonjones.org
There are at least two problems of course. I'm probably not that smart and I don't buy lottery tickets.
The article is more theory than practice. It practice, the big problem is not really about knowing what should be rebuilt. It is about tying the work of many people with different ideas on how thing should be done together.
The part about getting the build order correct, which is what the article is about is still important. But that's typically not the reason why everyone hates build systems. People can live with the occasional inefficiency, or "make clean". But getting that binary-only library that is not in the standard location with that source-only other library and its weird versioning scheme, both needing a different version of the same third library, written in a different programming language and using a different packet manager...
I personally hope Nix eats the world.
The paper is about thinking about build systems in a more modular way. Are dependencies specified statically or dynamically? What algorithm is used to decide what should get rebuilt after an update? What order do we rebuild things? What happens if we pick different combinations of algorithm+order?
For example, Make decides to rebuild based on timestamp, while Bazel decides to recompute based of file hashes. Thus, if you touch a file but don't change its contents, Make might rebuild things that Bazel wouldn't. Another example in the paper is the INDIRECT function in Excel. This makes it so that Excel cannot compute the dependency graph before the build starts. Excel has ways to deal with this, but they might cause some cells to be recomputed even when they didn't actually need to be. The paper has a bunch of interesting examples like this, it's worth the read :)
The problem of dynamic dependencies is particularly tricky. For example, suppose that you want to teach your makefile that your compilation depends on the source C file, but also on all the headers that are #included by that C file. And you want to actually look inside the C file to get that list of headers, not just copy paste the list into the makefile. There's more than one way to handle this.
For example, paper talks a great deal about schedulers, but those are very minute implementation details. If bazel gets "suspending" scheduler instead of "restarting" one, this is a super major change in the paper's eyes, but no user is likely to notice (other than "hey it works 2% faster now, cool!").
Moreover, the paper's scheduler abstraction is too limited - there are a whole bunch of schedulers which are non-representative in it, for example imagine a scheduler which is aware of memory and CPU usage, and schedules the tasks to avoid out-of-memory condition (like recent bazel).
The division on "cloud" vs "non-cloud" systems is also pretty weird... ok, bazel has cloud caching built-in but make requires external "ccache", but the resulting behaviors are pretty similar.
And the universal constant in all build systems is always a problem of management. I've turned 7 hour builds into 7 minute builds, and management doesn't care. I've saved literally thousands of hours of programmer time per year on projects, and management doesn't care. Build systems are a cost center. Always neglected, never funded, and not seen as critical infrastructure.
Build systems, at the end of the day, no matter the language or target platform, are never really that complicated, we just tend to over-complicate them. What I really want, if I had the time to implement, is a build file system.
All that said, I'll never go back to working on a big build system without having a PM that insulates me from the unreasonable demands of being woken up at 4AM by a Producer absolutely frantic because "the build system isn't working at all!" but when checking on the problem, it's because a developer checked in some code late and night and then went home. I have some serious trauma from those years of my life.