Mold Linker Version 3.0.0 Release – Rewritten in Rust
github.com
github.com
Now sadly we will have to fork and maintain the c version as mold2 forever.
Rust is not actually the right tool for all problems.
Do you have any reading material that you can share that might help me understand better? Docs for the distro, or an issue tracker I can search through?
In short, the entire distro is always built in one-shot at any given commit, and we can only rely on cached binaries from a past release if they or nothing in their supply chains changed. Given rustc depends on almost everything, mold3 would be built far too late to be useful for the most expensive build in the whole tree, which is rust.
Recursive dependencies would break our threat model, so we cannot use any rust tools to bootstrap rust. We bootstrap rust from llvm which we bootstrap from gcc which we bootstrap from tinycc which we bootstrap from M2Planet and so on back to 186 bytes of machine code.
Anyone getting a rust package from stagex must be able to build the entire tree up until that package and get the same hash, with no binary dependencies, thus removing any trust in maintainers.
If you can boostrap one version of stagex (so that its mold is trustworthy), I don't see why you can't use that to build another version of stagex.
As you say, you can rely on the cached binary because nothing in the old version's supply chain would change; it is pinned, and as long as you can build that, everything is fine.
In order for you to not have to trust us, you must clone our repo, of only source code, and build from zero to our released binary hashes. The shorter we can make the time for that to be possible, the more people we can convince to do it and ideally sign and publish their matching hashes. The more people we convince to do it, the less risk of us as maintainers being able add a backdoor without anyone noticing.
Mold was a tool to shave hours off the time most people have to spend doing a from-scratch verification. Forcing them to build a whole tree to get to rust to et to mold3, to then use that to build the whole tree a second time, would directly work against the goal of minimizing full tree verification time.
> Rust is not actually the right tool for all problems.
Rust is certainly the right tool for this problem, your own decisions notwithstanding.
Most popular Linux distros take a position of hoping and praying supply chain attacks do not target them. I am not convinced this will go well for them in the post AI world, but hey, I also hope I am wrong.
LLD would work fine for the whole tree, and did previously, but is much much slower than mold which is why we switched to mold.
Having to build all dependencies of rust including python and openssl and everything else before being able to use mold erases most of the full-tree build speedups as the path to rust is already about 80% of the full tree build time.
We are a distro that mandates independently verified 100% deterministic builds from source for any given release commit of the tree, so we have to build the whole tree several times for every release.
Would that meant that building rust is slower, but everything else is the same, and you don't have to maintain a fork of another complex project?
Rust is the single most expensive thing to bootstrap in any given linux distro.
It is the thing you need mold the most for to speed things up.
Unfortunately Rust does not even do full source bootstrapped deterministic builds themselves. Their release strategy is entirely the honor system, where everyone trusts their pinky swear that a single computer or person will never be compromised and allow for the injection of malware into the binaries everyone downloads from rustup.
The only reason we have even the crazy long bootstrap path we have today at all is the hard work of mutabah, an individual independent hacker that does it as a hobby.
The rust team seemingly considers memory safety as the only security problem that matters, and supply chain attacks out of scope.
Assuming the source code contains no issues (which can be checked later once the Rust compiler is built), one could leave that piece behind, and take the shortest path from .rs to executed code (C transpilation, or even an interpreter).
Then you can just use the latest built version to build mold and the next version of rust.
The problem is changing any dependency of rust, even python or perl or musl or openssl or the llvm stack or any dependencies of the llvm stack have the potential of resulting in a different hash for rust.
Any time a dependency is changed all decedents must be rebuilt. Which we must do very frequently. That is where mold saved us a ton of time.
Your stage1 will be based on whatever version is your stage0, but your stage2 should end up identical to any other build with any other stage0. So you don't have to rebuild the full chain all the time.
And then preferably you build a stage3 with FDO, you can easily get 20% more speed with a good corpus of tests.
Do you have at least a distributed cache with I guess sccache to save time for each rebuild?
Our goal is to allow people to reproduce the entire tree from source in the shortest time possible, so no one has to trust us, thus encouraging as many people to do it as possible, thus preventing us from having the means to inject supply chain attacks without anyone noticing.
It is literally a goal for new users to be able to run a server that bootstraps itself, and then bootstraps and signs all future releases, with remote attestation proofs. The shorter that initial build window is the better, which is why things like fast linkers written in early-bootstrappable-languages are so important.
> 186 bytes of machine code
This is better (smaller) than Forth!But I don't get why some people are obsessed with bootstrapping. Yes it is good to be able to do it, but it isn't something you need to do regularly.
Especially since rust had a much better cross compilation story than C or C++ (not as good as go or zig though), so you don't need to bootstrap on a new architecture, just cross compile to it. Furthermore, new architectures for hosting a compiler (as opposed to just a target, like microcontrollers) is a rare event. Just something that happens every few years.
It only matters if you have supply chain attacks in your threat model. Given they are up 400x since 2019, they should probably be in almost every threat model. Most distros operate on the honor system and that is not going to survive the post AI world.
Distros that are not fully hermetic and don't have reproducible packages (and there are many layers of reproducibility) will certainly have issues, but that's not a huge problem as long as you have a documented path to getting back to the current state. It doesn't need to be the fastest path, just a verifiable chain of trust.
It is critical to encourage many independent verifications that it be as fast as possible that someone can go from a clone of our tree of pure source code to the exact release hashes we publish.
Adding any binaries to that means someone that distrusts us must now go build those past releases as well, and if they rely on binaries, they must build those past past releases as well. This approach would make verification time go up dramatically every release.
Also, while you can use the latest prebuilt as a stage0, you could also use any other prebuilt from earlier down the chain as a stage0 and still get the same stage2 binaries. This is the next trust checkpoint you have, and it should be fine to have multiple ways to get there too, one for quick iterative releases and one that is reusing the minimal amount of bootstrapped packages.
Or you just try to only upgrade your stage0 once in a while when extremely necessary and it would build your new stage2.
So many ways to optimize the system, it's a choice to refuse to reuse what was trusted yesterday in order to build the next stage.
It's not a big deal for normal users, where you have Rust ready to go. Kind of a bummer in this case, but this is a specialized one.
This effectively triples or quadruples (maybe even more) the amount of code you need to trust for the cold bootstrap.
The fact remains that rust adds yet another toolchain to the boot process that needs trust and verification.
Unless the gcc rust engine is mature any time soon (lol), we have no path to use rust until very late game in a distro build.
The earliest we can bootstrap a go compiler is about 10 minutes. It builds directly from tinycc. Add 6 hours for our fastest compile of the shortest path to rust, with 192 cores.
I am a rust fan too, but it is the worst language to bootstrap, which is why for systems programming I still must often revert to C to have a small and reviewable and fast to build dependency surface.
(There is also a GCC backend called codegen_gcc that is pretty far along. And a separate reimplementation of both the frontend and backend using gcc and C++, called gccrs, which is not nearly as far along.)
> We will then conduct extensive compatibility testing and work closely with Linux distribution developers to make it practical for them to adopt mold as /usr/bin/ld. Making this happen is one of our highest priorities for mold 3.x.
"my problems (that are not mold's) are not solved by mold. How dare mold make those decisions?"
Maybe you should rewrite more of your linux distribution in rust so it's available earlier in the build process and get back to it being the default linker.
This is ridiculous. The amount of entitlement I'm reading here is gross. This is how you burn out maintainers and drive them away from open source.
> Your downstream should be precious.
No. Every open source maintainer is free to decide for themselves how much or how little they are willing to bend over backward for the sake of serving all possible user needs. If you don't like that, then feel free to build whatever you need yourself, from scratch.
Is my downstream paying me for the maintenance ?
If not, downstream may maintain mold2 themselves forever because their _extremely specific_ use case is neither a promise nor a valuable thing for mold to maintain. Debian understood this a while ago and isn't whining when they have to maintain their own fork. Maintainers maintain.
If enough downstreamers are unhappy about choices, they are also welcome to fork, until the base project is abandoned. Or they realize their usecases are extremely narrow.
(In addition, it's an extremely hypocritical and purist demand, because I am pretty certain that their distribution has, at some point, an arbitrary executable to make a compiler from. So they're probably okay with blobs, just not that one in particular.)
Yeah, for stuff we wanna rely on, this should ideally be the default. I'd go one step further and say no breaking changes past 1.0.0 at all. Instead people should favor creating entirely new projects (forks or not) and jump over to those, leaving the old one behind, if they want to do massive changes to something.
Of course, no one would be forced to do this, but it feels like if more did this, long-term supporting stuff that depends on those things would be a lot easier, if things could just be instead of changing under our feet all the time. Thank god for Nix and NixOS, even with their warts.
Edit: darn, it broke
I know this is common but it seems like either an aesthetic decision, or glibc cruft.
Bootstrapping at every build does not save you from the threat you think it does.
Using binaries from past releases is a strict downgrade in terms of verification speed, as it means a new independent reproducible build verifier must now build both trees, doubling the release verification time, and erasing any wins mold3 could otherwise offer.
Google can rely on lots of centralized internal provenance tooling to prove cached binaries are not tampered with to other Googlers but when the goal is proving end to end full source bootstrapped build integrity to any interested user from the public in the least time possible, the requirements are significantly higher.
Edit: this seems to have been cooking for a while when the first commit dropped: https://github.com/rui314/mold/commit/f41bfcd5c72ca30cce6498...
Yes. Comment by mold's author:
https://www.reddit.com/r/rust/comments/1w45j6n/comment/p7ac5...
(Wild is another fast linker, that only supports Linux)
Something I really look into because the crates situation (due to a too small stdlib à la R5RS Scheme) is Rust's Achilles' heel, since it directly affects its security claims.
This is true for - as an example - the WUFFS GIF decoder. You can get C which decodes GIFs and was transpiled from WUFFS, but that's awful code and nobody wants to modify that code, whereas the WUFFS source code for the decoder is fine.
When we look at rui314's changes to mold today after 3.0 release, they just modify the mold source code in Rust, as you'd expect if this was in fact now written in Rust.
Thus far, I've written two libraries that utilize Rust, and they are usually associated with "high performance". I was able to exceed the performance of similar libraries written in C/C++.
To be perfectly clear, this doesn't mean Rust can replace C/C++ in every scenario, but the actual perf gap is narrower than some folks may be willing to admit.
[0] https://shnatsel.medium.com/how-to-avoid-bounds-checks-in-ru...
In practice, projects written in Zig very much can choose both.
There’s also another caveat that there’s an effort to build in build time static memory safety checks in a way that’s more general than Rusts borrow checker, through a kind of plug in system rather than forced into the language. It’s not part of mainline Zig yet but this seems to be the direction Andrew wants to go.
Zig isn’t even in the same league as Rust regarding these things. Zig may still be around and active 10 years from now, Rust is guaranteed to be.
And second, did you use any AI tools for the rewrite?
The best programmers on the planet have tried and failed with this task for 50 years now, so I don't think this is true.
The main disadvantage of Rust right now is not supporting some more obscure platforms, but because mold wouldn't support them anyway I don't see that as a problem.
Last time I checked, I got impressed by the wide platform support, once you go down the tier list (https://doc.rust-lang.org/nightly/rustc/platform-support.htm...). What "obscure platform" specifically are you thinking about, that is currently missing from those lists?
LLVM doesn't have an equivalent annotated list, but it is missing alpha, bfin, c6x, fr30, frv, gcn, h8300, ia64 (aka Itanium), lm32, m32c, m32r, mcore, mep, microblaze, mmix, mn10300, moxie, nds32, nios2, pa (aka PA-RISC), pdp11, pru, rl78, rs6000, rx, sh (aka SuperH), storm16, v850, vax, and visium. For its part, LLVM does have some targets that GCC doesn't have (mostly related to GPU compilation).
The "big" targets that GCC has that LLVM lacks are Alpha, Itanium, PA-RISC, and SuperH, with Itanium being sufficiently weird that it's pretty firmly in the "fuck this" category from a maintainer's perspective, and people are actively ripping out support for it.
However, there is now a Rust codegen plugin for GCC, so even this disadvantage is now basically moot.
Rewrite all the things!
There are many. Two big ones are Alpha and PA-RISC. NetBSD and Linux continue to support both. Linux distro choices are pretty much limited to Gentoo though.
2. LLVM does not support as many backends as GCC does, so even if you did get 100% of the LLVM supported backends up and running, you'd still be missing some.
But even for the architecture part there can be differences. For example, my understanding is that a lot of the calling convention details need to be handled by the frontend in LLVM, leading rustc to duplicate logic from clang here. Things like varargs FFI with C code can be particularly gnarly.
IMO, given the recent commits: almost certainly.
The only way I found that works reliably is stick to a small well defined set of mostly safe primitives. E.g. at work we use message passing / event bus architecture everywhere, which works great for a robotics / industrial context. But even then, if you somehow mess up and have a variable accessed from event handlers in two different threads, it is tough to spot other than if you get lucky and observe it with a build using TSAN.
With rust that class of mistakes is just entirely eliminated, which makes it easier to to concentrate on the hard things that actually matter (like the domain specific logic).
It's like "in a Zodiac boat / aircraft carrier navy". This customary putting C and C++ into the same bucket is as amusing as it is unproductive.