But I still do not understand the point of "reproducible builds". I know what they are, but to me the amount of work involved outweighs the benefit.
I even heard NetBSD is also working on "reproducible builds". So maybe I am missing something :)
But I still do not understand the point of "reproducible builds". I know what they are, but to me the amount of work involved outweighs the benefit.
I even heard NetBSD is also working on "reproducible builds". So maybe I am missing something :)
The main benefit is that you can trust that the resulting binary file being served matches the source code that it's build from. This mostly matters for distros in that they build from a source package repository, but anyone running a mirror could hypothetically replace the package with another (potentially malicious) package, leading users to install malicious tooling. It mainly matters for distros because pretty much every distro out there runs on third party mirrors (often ran by universities, but also just people who want to help) rather than on direct upstream; packages get uploaded to a main server, then mirrors copy from that main server (to reduce network traffic load on the main server). Right now, mirror trust is mostly "we assume you're not gonna be evil, until we get complaints". If the build is reproducible, the software can inherently confirm that the file they're getting is trustworthy, making "getting complaints" much easier to confirm.
It can also speed up the overall building process; if the package source code hasn't changed, you can also always assume that the resulting binary hasn't changed (meaning you can use hashes instead of relying on mtime like make does). Docker build cache works in a somewhat similar way (although docker isn't inherently deterministic).
Devwise, you can also reconstruct a build much easier if it's reproducible; ie. if you've accidentally thrown away the .elf file for debugging, if your build is deterministic, you can just rerun the build and get the same .elf file again.
[0]: While not a problem for Linux distros, in cases where you need a secret to sign an application, reproducible typically means "identical except for the signature" instead. F-Droid uses this for example to figure out if they should use buildserver stuff or the original APKs: https://f-droid.org/docs/Reproducible_Builds/
It was my assumption that a mirror is required to host a build that has a hash conforming to the original. Is that not the case?
I thought all packages were cryptographically signed, and that the package manager would compare the hashes of artifacts downloaded from mirrors to the hashes listed in the package index (which is also signed). This is not an attack that needs reproducible builds to mitigate.
Note that mtime still has the advantage of being faster than hashing.
I think the actual main utility is that the process has done a very good job of rooting out several causes of unintentional nondeterminism in the build process. I say unintentional because the two main causes of unreproducibility, by several orders of magnitude, are timestamps being embedded everywhere and absolute paths being embedded everywhere, and those are rather expected. But some of the unreproducibility comes from things like accidental reliance on inodes in file paths (i.e., doing "for file in listdir()" without sorting the results of listdir) or the compiler itself accidentally sorting based on pointer address (which is unreproducible on ASLR systems).
I don’t know whether I’d spend this much work on such an abstract goal, but what reproducibility changes really is quite amazing. It vastly increases trust in published binaries and obviates the need for signing and the security benefit of compiling software yourself.
Not really. Most people still would rely on signatures because they can't be expected to compile everything from scratch just to verify their download is authentic. Moreover even though reproducible builds make verification easier, it still requires someone to sound the alarm. For less popular packages there might be nobody checking any particular build is backdoored, because most people see "reproducible builds" and they assume Somebody Else is doing the reproduction.
you still may not trust the public gobuilds instance. my hope is that people (eg software projects themselves, or distros, or other kinds of communities) will run & use their own gobuild instances and verify their builds against the public gobuilds service. win-win: gives them assurance their builds are really reproducible, and builds trust in the public gobuilds (keeping it honest, if someone sees a hash mismatch, they will speak up).
i usually don't get much enthusiasm for it though. (:
Given the increasing likelihood of supply chain attacks, isn’t this a very prudent precaution?
If I'm understanding correctly, the malicious code was introduced as part of the test code, so no matter who compiled it, they'd get a binary with the same (malicious) functionality. Heck, it might even have been reproducibly malicious.
The real crazy part was that it was modifying the functionality of sshd at runtime, allowing the attacker to log into any system.
Reproducibility of either sshd or xz wouldn't have stopped this attack.
That's my reading of https://research.swtch.com/xz-script.
Yeah, `bazel run test:...` would have access to the test files, but `bazel build xz:executable` would not (by default) be able to pull in extra shenanigans from the test files (and I think there's generally linting and formatting rules required by default with `BUILD.bazel` files, reducing another sneak-vectors)
For that we'd need some sort of source code reviewing effort like https://github.com/crev-dev/cargo-crev implements. I've started whatsrc.org to keep track of the source code inputs we're putting into our computers (that would benefit from reviews), but the conclusion is also somewhat "it's too much".