Build Systems à La Carte: Theory and Practice (2020) [pdf]
microsoft.com
microsoft.com
also, earlier:
Build systems à la carte: theory and practice (revised and expanded) [pdf] - https://news.ycombinator.com/item?id=25113759 - Nov 2020 (3 comments)
This paper guided the design. It's one of my favorites, and I read it once a year.
Once it is, these will be possible. And I did misspeak: filesystem watching is not easily done by dynamic dependencies, but it will be easy for Rig.
Speaking as though Rig is full-featured, tracing command dependencies can easily be implemented by users (by adding strace on commands), and then the data from those traces can be stored in the build database. Then dynamic dependencies can be used to make those as dependencies on subsequent builds.
For filesystem watchers, Rig has the capability of having different build types. One of the functions of a build type is how to check a file for modification.
As of right now, it just returns true. Once incremental builds are added, files will be checked for modification in that build type's function.
To add support for filesystem watchers, I simply add a build type that checks the watcher for modifications in that function.
And for command tracing, the most significant issue is that it records both files read and files output. Even with dynamic dependencies it is difficult to record both of these dependency types properly. For example, Shake has this "multiple outputs" feature but it is actually really hacky and doesn't work. https://github.com/ndmitchell/shake/issues/382 The Pluto "sound and optimal" build paper discusses this briefly - they use a two-level file-command graph specifically to solve this problem. Since you have files-as-targets you will have trouble recording write dependencies, similarly to Shake.
Okay, but I don't really consider that a problem.
Computers are fast enough that burning those cycles, even for 1 million C files will take less than a second, probably less time than a user would notice. One hundred ms would be my target, and that means I would have about 100 cycles per file, plenty of time if my data structures are good enough since all I am doing is setting a status enum.
People would notice if Rig stat'ed all of the files, so I still want a watcher. But people forget how fast computers are, and I'm not about to complexify Rig to be faster when the user is not likely to notice.
> And for command tracing, the most significant issue is that it records both files read and files output. Even with dynamic dependencies it is difficult to record both of these dependency types properly.
Yeah, and? This is easily done with the right design. Rig has multiple outputs built in. You can create targets as you're executing other targets. You will be able to add a dependency both ways, by adding a dependency to the currently executing target, or adding the currently executing target (or another target) as a dependency of another target.
It is "difficult to record both of these dependencies properly" because build systems are built with limits. Even Shake.
Users do notice, one of the comments about Tup is that it is surprisingly fast. Now that could be other things like being written in C and using sqlite but I'd imagine at least some of it is due to being designed to use efficient algorithms.
> build systems are built with limits
But the limits are self-imposed. It is not too hard to extend a build system to add these features, but even before Rig is fully done, you seem to be in "maintenance mode" where no new features are allowed. I myself have also been working on designing a build system (and programming language), but I have deliberately avoided writing any code so that I am not tied down by refactoring when I discover a new paper or approach. At this point though my literature survey is complete, so it is time to start coding. I was thinking perhaps you beat me to an implementation, but I guess not.