Disorderfs: FUSE-based filesystem that introduces non-determinism into metadata
salsa.debian.org
salsa.debian.org
As I understand it, currently several packages are not reproducible (like, python, pytest, gcc), so it is not priority, but when those large packages will be done, r13y will start using DisorderFS to uncover remaining reproducibility bugs.
This is too idealistic, but gives lots of pleasure about package space.
Once such criterion is enforced, then everybody can reproduce the build with that set of source files and build instructions, without requiring a special environment that forces a specific order of events.
...by using non-determinism?
Very mind bending for me, I'm not sure I understand, but I'm glad smart people are figuring this stuff out.
(At least this is how Debian uses disorderfs. I wrote the first version of disorderfs 6 years ago in a hacking session at DebConf15 in Heidelberg. I never expected to see it on the front page of HN!)
Therefore (the original question), instead of using "disorderfs", why not write and use an "orderedfs" for every build?
And since this is run as part of a CI process, you will get lots of builds over time and will root out all sorts of issues caused by non-determinism.
As to your original question, there are so many sources of nondeterminism that trying to emulate them all away would make builds more complicated, less performant (FUSE adds overhead), and less safe (since there would be more components that could potentially be backdoored).
https://tests.reproducible-builds.org/debian/index_variation...
It is similar to how Chaos Monkey increases the resilience of Netflix's service by introducing random failures and then for each of those failures working out how to prevent the failure from affecting the overall status of the service.
That would only fix the build machine's problem, it wouldn't fix anyone else's builds.
A repeatable build without determinism is a fix for all people everywhere.
It would fix everyone else's builds if they added determinism in their builds (as opposed to adding non-determinism in the test-procedure). Besides, a test-procedure never gives a guarantee because a bug depending on non-determinism can be subtle.
That's the point - to uncover bugs dependent on nondeterminism by using a filesystem that introduces it. This is for fuzz testing at the filesystem level, not literally reproducing the builds correctly multiple times.
From the linked README:
"This is useful for detecting non-determinism in the build process."
https://reproducible-builds.org/citests/
All distros there are essentially reproducible fuzzing CI systems that introduces determinism through disorderfs, lang changes and so on. These changes are fine to fix but nothing you'd normally get when reproducing packages for a distribution.
Personally the important part is if the patches are upstreamed or not. This isn't something that is a priority among distribution.
Results from a fuzzing in Arch:
https://tests.reproducible-builds.org/archlinux/archlinux.ht...
Results from just chroot recreation:
https://github.com/buildbarn/bb-remote-execution/blob/eb1150...
Docker instead has to re-build all subsequent layers when the input of even just one layer changes.
The two tools sit on a different point in the spectrum of simplicity of use though. Maintain build files for bazel (and dealing with other constraints of hermetic execution) is hard time-consuming and it's hard to convince many teams it's worth the effort.
Docker apparently struck a sweet spot in that it's easy to explain how to craft linear build steps and it does a half-decent job in actually caching stuff. Sometimes it doesn't work well depending on your workflow but people then have the incentive to read about how to improve the "cacheability" of their dockerfiles (multistage, reorder, dockerignore, ...). As many things in our craft, the human aspect trumps over technical brilliance.
Alas that's not what this is about, but now I wonder how hard it would be to make the thing I had in my head.
https://blog.nixbuild.net/posts/2021-01-13-finding-non-deter...
It isn't enabled by default but can be turned on with a setting: