Reproducibility is technically pointless, because you still have to trust the developer, and they can still add backdoors.
Reproducibility is technically pointless, because you still have to trust the developer, and they can still add backdoors.
Builder != developer - and with reproducible builds, you no longer beed to trust the builder. CI is commonly used for the final distributable builds and you can't always trust the CI server. Even if you do, many rely on third party thingd like docker images - if the base build image gets compromised, code could trivially be injected into builds running on it and without reproducible builds, that would not be detectable.
As a developer, it would be quite reassuring to build my binary (which I already do for testing) and compare the hash with the one from the CI server to confirm nothing has been tampered with. As a bonus, distro maintainers who have their own CI can also check against my hashes to verify their build systems aren't doing something fishy (malicious or otherwise).
That makes sense! However, this is not a good argument for reproducible builds, because you can already do that today.
You already have to build a trusted binary locally for testing right? You're dreaming of being able to compare that against the untrusted binary so that you can make sure it's a trusted binary too - but you already have a trusted binary!
Okay - but it's a hassle, you don't want to have to do that, right? Too bad - reproducible builds only work if someone reproduces them. You're still going to have to replicate it somewhere you trust, so you gained practically nothing.
You can also have ten people on the internet verify the untrusted binary. With signatures, adding more people doesn't help.
That's not how it works, you have to reproduce it before it becomes trusted.
> You can also have ten people on the internet verify the untrusted binary.
Sure, then we have to build a complex consensus system that introduces a bunch of unsolved problems. My opinion is that this just isn't worth it, there is practically nothing to gain and it's really really hard.
Eh, there's stuff you can do with software before you trust it. Eg you can start pressing the CDs or distributing the data to your servers. Just don't execute it, yet.
> Sure, then we have to build a complex consensus system that introduces a bunch of unsolved problems. My opinion is that this just isn't worth it, there is practically nothing to gain and it's really really hard.
It's the same informal system that keeps eg debian or the Linux kernel secure currently:
People don't do kernel reviews themselves. They just use the official kernel, and when someone finds a bug (or spots otherwise bad code), they notify the community.
Similar with reproducible builds: most normal people will just use the builds from their distro's server, but independent people can do 'reviews' by running builds.
If ever a build doesn't reproduce, that'll be a loud failure. People will complain and investigate.
Reproducible builds in this scenario don't protect you from untrusted code upfront, but they make sure you'll know when you have been attacked.
There's a big difference here. When a vulnerability is found in the Linux kernel, that doesn't mean that you were compromised.
If a build was found to be malicious, then you definitely were compromised and it's little solace that it was discovered after the fact. This is why package managers check the deb/rpm signature before installing the software, not after.
This is just an additional check that the debian repository has sane builds.
(If someone mucks around with the debian repositories but you aren't the target, you might or might not be under attack.)
A talented developer might still be able to create a bugdoor which gets past code review, but that takes more effort and skill than just putting the malicious code into a local checkout and then saying "How did that get there?".
You can already verify that a toolchain wasn't backdoored today, reproducible builds aren't necessary for that.
How, exactly?
If we both compiled hello.c (a prototypical hello world program), and exchanged binaries; how would you verify my build wasn't malicious?
That does require reproducible builds, but here is how to do it without reproducible builds:
Take the trusted source code, then compile it to make a trusted binary. Now put the untrusted binary in the trash, cause you already have a trusted binary :)
You are obviously familiar with Bazel/Blaze etc. Wouldn't reproducibility be necessary for those systems to work well most of the time? I can think of exceptions (like PGO), but it seems useful to produce at least some binaries this way. Also covered in this: https://security.googleblog.com/2021/06/introducing-slsa-end...
That depends, I think it's difficult and mostly still pointless. I wrote about this a bit in the blog post I linked to. It's a big trade off, for questionable benefit.
> Wouldn't reproducibility be necessary for those systems to work well most of the time?
Yes, there are definitely some good non-security reasons to want deterministic builds. My gripe is only with the security arguments, like claims it can reduce threats of violence against developers (!?!).