Do you have any reading material that you can share that might help me understand better? Docs for the distro, or an issue tracker I can search through?
Do you have any reading material that you can share that might help me understand better? Docs for the distro, or an issue tracker I can search through?
In short, the entire distro is always built in one-shot at any given commit, and we can only rely on cached binaries from a past release if they or nothing in their supply chains changed. Given rustc depends on almost everything, mold3 would be built far too late to be useful for the most expensive build in the whole tree, which is rust.
Recursive dependencies would break our threat model, so we cannot use any rust tools to bootstrap rust. We bootstrap rust from llvm which we bootstrap from gcc which we bootstrap from tinycc which we bootstrap from M2Planet and so on back to 186 bytes of machine code.
Anyone getting a rust package from stagex must be able to build the entire tree up until that package and get the same hash, with no binary dependencies, thus removing any trust in maintainers.
If you can boostrap one version of stagex (so that its mold is trustworthy), I don't see why you can't use that to build another version of stagex.
As you say, you can rely on the cached binary because nothing in the old version's supply chain would change; it is pinned, and as long as you can build that, everything is fine.
In order for you to not have to trust us, you must clone our repo, of only source code, and build from zero to our released binary hashes. The shorter we can make the time for that to be possible, the more people we can convince to do it and ideally sign and publish their matching hashes. The more people we convince to do it, the less risk of us as maintainers being able add a backdoor without anyone noticing.
Mold was a tool to shave hours off the time most people have to spend doing a from-scratch verification. Forcing them to build a whole tree to get to rust to et to mold3, to then use that to build the whole tree a second time, would directly work against the goal of minimizing full tree verification time.
> rust from llvm which we bootstrap from gcc which we bootstrap from tinycc which we bootstrap from M2Planet and so on back to 186 bytes of machine code.
While the effort is laudable, the fact that these are built from source is not enabling me to verify them. Rust, LLVM, gcc, tinycc, all the way back to that 186 seed are prohibitively large to verify.
I'm not even in a position to verify the delta to these projects over a single day. I'm having to trust _someone_, many someones.
Our specific responsibility is to limit the number of people users have to trust to only the actual authors of the software they are installing.
> Rust is not actually the right tool for all problems.
Rust is certainly the right tool for this problem, your own decisions notwithstanding.
Most popular Linux distros take a position of hoping and praying supply chain attacks do not target them. I am not convinced this will go well for them in the post AI world, but hey, I also hope I am wrong.
Okay, but what if all the signers are people that a new user does not know? Who is to say all those signatures are not a bunch of made up AI identities?
This is why we make it trivial for anyone to clone our tree and build from zero and get the same result at any release with no required binaries of any kind anyone has to verify the provenance of. This is called full source bootstrapping. The cheaper we make that, the more people that will do it and the higher the chances users will see a signature from someone they personally trust.
I’ve been told “no one thinks a programming language always makes software safer” but the comments here suggest many actually do believe just that. The problem you’ve defined is a threat to the “my language is a silver bullet” belief.
Often having to make lots of small patches to preserve upstream functionality while fixing their obvious security and determinism bugs. Or we have to ignore autogenned code and figure out bootstrapping ourselves. Many upstreams just put binaries in their source code and call it reproducible.
For instance, XZ published a malicious hand-packed archive of their code that many distros use because it has all the auto-generated code already. We totally ignored that archive and pulled the code that was actually reviewed, and ran autogen ourselves. In doing so we were never impacted by the XZ attack.
We are obligated to do anything we can to protect our users, even if most thing doing so is paranoid.
How about a rust-to-wasm process, you store the wasm binary, but have a human readable wasm interpreter that only implements enough for mold to work? Would that be trusty enough?