If you mean which distribution has 100% of its packages reproducible, probably none yet. But Arch and Debian are both making progress.
> 56 (100.0%) out of 56 built NetBSD files were reproducible
That is not really comparable to debian's 26475/28522 for buster https://tests.reproducible-builds.org/debian/reproducible.ht...
OS-wise, NixOS.
Build system-wise, there are lots of options: Blaze, Buck, Pants, Please (AFAIK)
In case any Nixers are reading this, here is how I got NSPR to build reproducibly in Guix:
https://git.savannah.gnu.org/cgit/guix.git/commit/?id=6d7786...
https://r13y.com/ tracks the progress of NixOS reproducibility; currently we're at 98.23% bit-for-bit identical for our minimal installer ISO. After that, we'll need the graphical installer, and then more of the base package set. So we've still got a ways to go.
If you build the same code on two different machines, using the same compiler, with the same options, then the generated binaries should be exactly the same.
A build process that names things with timestamps or leaks your locale into the build configuration (or doesn't pin build-time dependency versions) will make the build depend on things other than the source code (both program and build settings) you made available.
It may even be desirable for it to be non-reproductible - if, for instance, you want to use optimizations targeted to your specific system, then your build system will have to introduce the architecture information into the build process and your build will result in a unique binary that targets your own machine.
For example, depending on the input order, linker may produce different output. Surely you can sort the object files, but the sorted object files order is still effectively "stored" into the binary, and that's not source code.
You can only normalize such things (like in the example above, sorting), you can not eliminate them, they naturally exist.
No, but the order should be explicitly defined in the build scripts or the result will not be deterministic.
If the order triggers, say, a linker bug that makes one in 50 builds crash, execution will not be deterministic and that's really, really bad.
There is so much context that is normally embedded into a binary that this is usually not true unless explicit measures have been taken.
Two very common sources that introduce variability are time-stamps used in the build, and environment variables such as $HOME and $USER.
That's 100% controllable and deterministic.
Let's assume there is a latent bug in the compiler that gets triggered if file four is the first one. Good luck debugging that.
Most compilers give no guarantees in which order they lay out the data. I love deterministic processes as much as everyone. But randomized approaches have their advantages too. And if a compiler has reasons to randomize output e.g. for speed than it’s a trade off to consider.
grabs lock
writes to file
writes to index
releases lock
That's not a race condition. The output order doesn't matter, but it is nondeterministic.
It's not true in practical code either, people like to stick in timestamps.
It's not ever true on windows, unless you use the fairly recent PE header changes.
So I can checkout an arbitrary version from years ago and reproduce the exact same set of output files?
Think of it as absent a cache I should get bit for bit identical out (perhaps ignoring logs and such).
I maintain my own build farm and tried comparing my results against the official CI server:
$ guix challenge --substitute-urls="https://ci.guix.info"
14,224 store items were analyzed:
- 4,972 (35.0%) were identical
- 265 (1.9%) differed
- 8,987 (63.2%) were inconclusive
Of the 5237 build artifacts that were available on the substitute server, only 265 (5%) differed.All of these items can be (and have been) built entirely from source, starting with Guix' initial "binary seeds", on (probably) different hardware and kernel compared to the CI system.
One reason builds become irreproducible is when a build is multi-threaded, and the order in which artifacts are combined into larger ones becomes unpredictable. That problem doesn’t exist, or at least is a lot smaller, for ‘leaf’ artifacts (example: if your C compiler is single-threaded, and you run make multi-threaded, individual object files do not have the ordering problem, but libraries built from multiple object files do)
On the other hand, a single static struct with a padding “hole” that isn’t consistently written that happens to end up in lots of binaries will decrease your percentage a lot.
Each of these "artifacts" are actual isolated builds of complicated programs such as Chromium or GCC. The technical term is "derivation", which produce "outputs".
All of those packages can be reproduced from source now or 100 years into the future and SHOULD produce the exact same binary output. If they don't, it's a bug.