I (naively, apparently) assumed this had been possible with open-source toolchains for a long time.
I (naively, apparently) assumed this had been possible with open-source toolchains for a long time.
Reproducible builds avoid all this and always produce the same outputs given the same inputs. There's no good reason (that I can think of) why this shouldn't have been the case all along, but for a long time I guess it just wasn't seen as a priority.
The benefit of reproducible builds is that it's possible to verify that a distributed binary was definitely compiled from known source files and hasn't been tampered with, because you can recompile the program yourself and check that the result matches the binary distribution.
Well, it's not like developers consciously thought "How can I make my build process as non-deterministic as possible?", it's just that by the time people started to become aware of the benefits of reproducibility, various forms of non-determinism had already crept in.
For example, someone writing an archiving tool would be completely right to think it is a useful feature to store the creation date of the archive in the archive's metadata. The idea that a user might want to force this value to instead be some fixed constant would only occur to someone later when they noticed that their packages were non-reproducible because of this.
But you're right; if the goal had been thought of from the start, there's no reason why every build tool wouldn't have supported this.
It's not just security. If a hash of the input sources maps directly to a hash of the output binaries, then you can automatically cache build artefacts by hash tag and get huge speedups when compiling stuff from scratch.
This was the primary motivation for Nix, since Nix does a whole lot of building from scratch and caching.
---
There is Debian initiative to create bit-to-bit reproducible builds for all their software (well, all critical).
https://reproducible-builds.org/
R13y is akin to "computer proofs" in math -- if you don't have it, that's fine, but if you have it, that's awesome.
There are practical reasons to favor reproducibility too, but those are more for distro maintainers.
The fact that NixOS (not Debian) got this 100% is mostly because
- minimal image has a small subset of packages (https://hydra.nixos.org/build/146009592#tabs-build-deps)
- Nix tooling was created 15 years ago *exactly* for this, Nix is mad to make packages bit-to-bit rebuildable from scratch.
- Nix/Nixpkgs is growing in number of maintainers and got more funds
- Nix has fewer Docker/Snap pragmatics
I don't think this is accurate?
Nix is about reproducing system behaviour, largely by capturing the dependency graph and replaying the build. But this doesn't entail bit-for-bit identical binaries. It's very much sits in the same group such as Docker and similar technologies. This is also how I read the original thesis from Eelco[0].
And well, claims like this always rubs me the wrong way since nixos only really started using the word "reproducible builds" after Debian started their efforts in 2015-2016[1], and started their reproducible builds effort later. It also muddies the language since people are now talking about "reproducible builds" in terms of system behavior as well as bit-for-bit identical builds. The result has been that people talk about "verifiable builds" instead.
Determinism can decrease performance dramatically. Like concatenating items (say, object files into a library) in order is clearly more expensive in both time & space than processing them out of order. One requires you to store everything in memory and then sort them before you start doing any work, whereas the other one lets you do your work in a streaming fashion. Enforcing determinism can turn an O(1)-space/O(n)-time algorithm into an O(n)-space/O(n log n)-time one, increasing latency and decreasing throughput. You wouldn't take a performance hit like that without a good reason to justify it.
Presumably, incremental compilation is only for development. For release, you would do a clean build, which would be reproducible.
> Even just stripping your compile paths from debug info is not entirely straightforward
Just use the same paths.
I’d say that’s exactly the wrong approach: given how hard incremental anything is, it would make sense to insist on bit-exact output and then fuzz the everliving crap out of it until bit-exactness was reached. (The GCC maintainers do not agree.) But yes, you could do that. It’s not impossible to do reproducible builds with GCC 4.7 or whatever, it’s just intensely unpleasant, especially as a distro maintainer faced with yet another bespoke build system. (Saying that with all the self-awareness of a person making their own build system.)
> Just use the same paths.
I mean, sure, but then you have to build and debug in a chroot and waste half a day of your life figuring out how to do that and just generally feel stupid. And your debug info is still useless to anybody not using the exact same setup. Can’t we just embed relative paths instead, or even arbitrary prefixes it the code is coming from more than one place? In recent GCC versions we can, just chuck the right incantation into CPPFLAGS and you’re golden.
All of this is not really difficult except insofar as getting a large and complicated program to do anything is difficult. (Stares in the direction of the 17-year-old Firefox bug for XDG basedir support.) That’s why I said it wasn’t a GCC problem so much as a maintainer attitude problem.
By being able to reproduce the file completely, down to identical md5 hashes, you know you have the same file the creator has, and know with certainty that the file has not been tampered with
Concretely, you would need to keep track of and reproduce e.g. the march flag value as a part of your build input. If you wanted to optimize for multiple architectures, that would mean separate builds or a larger binary with function multi-versioning.
If you want to compile any piece of software available in Nixpkgs, you can override it's attributes (inputs used to build it).
One can trivially have an almost identical operation system to your colleagues install, but override just one package to enable optimisations for a certain cpu. This would however imply that you'd lose the transparent binary cache that you could otherwise use.
Exactly this method is used to configure the entire operating install! Your OS install is just another package that has some custom inputs set.