Running the "Reflections on Trusting Trust" Compiler
research.swtch.com
research.swtch.com
But I guess, I'm too used to it, too?
If you follow the Wiki though, I think you still learn quite a bit: https://wiki.archlinux.org/title/installation_guide
I started with Arch around 2010, and last did a fresh install on a new computer a few months ago.
Yes, you can still learn quite a bit from the guide if you are new to it. But it's gotten a lot easier over the years. I remember eg X and wifi used to be a bit of a pain, but now they mostly just work. Thanks to UEFI, you can also boot directly into your kernel, and no longer need to muck around with an extra boot manager like grub or syslinux. (Assuming your computer's UEFI implementation is not too buggy.)
For a job, we needed Go on FreeBSD 8, which entailed a local patch to revert pipe2 to pipe in Go, and building from source. That was my first bootstrap. Go, as Russ mentions, is easy to bootstrap, and I've done it a couple of other times.
I've tried to bootstrap Rust, from its old compiler written in OCaml, but that one is trickier. It has not been maintained like Go's bootstrap compiler and the bootstrap chain has long since been broken. Furthermore, rustc is only guaranteed to be able to be buildable with the previous release. As far as I can tell, no one has built rustboot in many years. I like doing software archaeology, so I'll probably try again with that project sometime.
mrustc takes another approach, closer to the aforementioned GNU Mes, in that it's a reimplementation of Rust, intended for bootstrapping rustc. https://github.com/thepowersgang/mrustc
- git clone itself in a subfolder
- git checkout and build the initial D compiler
- install it in a temporary prefix
- git checkout the first breaking commit and build it with the initial compiler
- install it over the previous compiler
- git checkout ...
and so on. In normal operation all these stages would be cached, so before I abandoned this approach, I think I was up to a hundred or so intermediate versions.
Eventually I started doing git releases, which needed a better solution. So since I have an optional C backend, I just build the compiler with the C backend, then zip up all the C files to make the release. Then to bootstrap from it, I just do (effectively) gcc *.c -o build/neat_bootstrap.
edit: Ah, here it is: https://github.com/Neat-Lang/neat/blob/v0.1.6/bootstrap.sh
Rust has a long chain too, but rustc only uses features itself, that the previous release supports. Releases are every 6 weeks.
Go had been bootstrapping from 1.4 (the last C compiler release), until the release of 1.20 this year, when the bootstrap compiler was bumped to 1.17.13 and will be bumped yearly [0]. That meant go1.4 had to be able to compile new versions, keeping the compiler from using new features in itself for 8 years. Notably, this now allows for generics to be used in the compiler.
https://bootstrappable.org/ https://github.com/oriansj/stage0 https://github.com/oriansj/stage0-posix/
The difficulty of bootstrapping GHC
Second, and more importantly, many people carry a copy of a trusted compiler around, though it’s rarely mention in attacks like these: their head. In a pinch people can do spot checks to verify codegen to see whether it looks correct, unless the backdoor is incredibly subtle. But experience shows us that the more complex and hidden a backdoor is the more likely it is to break when subjected to unfamiliar examination.
> In broad strokes they do but often they’ll have randomness creep in, such as different build hashes or iteration order of associative containers.
In order to find defects related to codegen being altered by incidental ordering of objects in containers, a build option called LLVM_ENABLE_REVERSE_ITERATION was created. Builders periodically run with this mode to verify that the regression test suite still passes when containers iterate in reverse.
That said, it's true that there are probably remaining sources of variation among builds. There is a significant effort [1] to find these and avoid them.
I would guess that this may also end up favoring the execution performance of the compiler over other design choices, but I could be wrong about that one.
Timestamps are one of the biggest offenders, but this is also why reproducible builds are important. Nondeterministic codegen is just scary.
Second, and more importantly, many people carry a copy of a trusted compiler around, though it’s rarely mention in attacks like these: their head.
This is also why I'm against inefficient bloated software in general: the bigger a binary is, the easier it is to hide something in it.
Along the same lines, a third idea I have for defending against such attacks is better decompilers --- ideally, repeatedly decompiling and recompiling should converge to a fixed point, whereas a backdoor in a compiler should cause it to decompile into source that's noticeably divergent.
>Nondeterministic codegen is just scary.
Why not focus on deterministic codegen only? Caring about some metadata changing does not seem to be as useful from a security perspective.
If the metadata were neatly separated from the executable portions of an artifact then it indeed wouldn't be too much of an issue[1], but it often isn't. It sometimes gets embedded into strings and other places in the executable itself where it theoretically could have an impact on the runtime (and of course makes executable signatures unreliable, etc.). Without that separation comparing two potentially-the-same executables becomes equivalent to the Halting Problem.
[1] You could sign only portions of the executable, for example... but such signature schemes are brittle and error-prone in practice, so best avoided. So you'd ideally want any and all metadata totally separate from the executable.
If there is a backdoor in the code there is a backdoor regardless of if there is an extra timestamp that exists.
This way anyone who signs modified code can be caught. This is especially easy with open source projects. Incidentally, Debian is working very hard on deterministic builds.
Code signatures have already existed for decades and do not require reproducable builds.
Without deterministic builds, it's possible for the attacker to switch the compiled binary release and sign it with a stolen key with nobody noticing.
I'd argue that if the key being stolen/misused becomes more easily detectable, signing with said key becomes more meaningful.
Certificate Transparency logs only show when new certificates are issued. If a certificate is stolen it is of no use.
It lets issuers show they are trustworthy.
If the NSA could steal a debian key (or coerce a packager), without deterministic builds, they could put a backdoor in a popular Debian package with barely any chance of getting caught.
Deterministic builds make it much easier to detect this. That makes the value of such a stolen/coerced key much lower. Making it much less likely this gets attempted, and likely stopping anyone who used to try this attack.
If you have to make exceptions for some metadata, all of a sudden there's judgement and smarts involved. And judgement can go wrong.
>And judgement can go wrong.
Which is why it is important to have a good design and threat model.
That page has progress reports that are 8.5 years old and the project is not close to be done. There are still thousands of Debian packages to go.
However, at least the problems with this approach are all in one direction: Bits differing might be due to an actual attack or because of weird compiler issues, but if you get the bits to be identical, you know that no funny business is going on.
The proposal to have more complicated rules doesn't have this upside. And you still need to make sure that the things that actually matter are the same (like instructions in your binaries), even if you successfully separated out irrelevant metadata changes.
but why would you assume that the decompiler is not backdoored? It would know to remove the backdoor code, so the fixpoint is not going to show anything.
In practice, you can push the likelihood of this down to arbitrarily small levels. Pick a decompiler that's as unrelated to your compiler as you can find. Then pick a second decompiler that's maximally unrelated to both the compiler and the first decompiler. It's highly unlikely all three tools will be backdoored in a fully compatible way.
This is the whole point of Thompson’s lecture, though.
Not exactly easy, but probably easier than a decompiler that produces human-equivalent source code.
It could be an interesting exercise to bootstrap up from something like this to a working linux environment based solely on source code compilation : no binary inputs. Of course a full linux environment has way too much source code for one person or team to audit, but at least it rules out RoTT style binary compiler contamination.
By default, but https://reproducible-builds.org/ has made excellent headway on allowing that to be fixed if you care.
From the article:
> Today, however, the Go compiler does compile itelf, and that prompts the important question of why it should be trusted, especially when a backdoor is so easy to add. The answer is that we have never required that the compiler rebuild itself. Instead the compiler always builds from an earlier released version of the compiler. This way, anyone can reproduce the current binaries by starting with Go 1.4 (written in C), using Go 1.4 to compile Go 1.5, Go 1.5 to compile Go 1.6, and so on. There is no point in the cycle where the compiler is required to compile itself, so there is no place for a binary-only backdoor to hide.
I think that this paragraph gives just hints on where to look for:
* there is no "binary-only backdoor", so I'm expecting the backdoor to be available in the source distribution of Go.
* the backdoor is either in each release and/or is propagated from one release to the other
There is a feature of the Go compiler that I mistrust: //go:embed. I expect it could be a good vehicle to inject source code that isn't in the repository as a file. Especially as Russ himself is deeply involved in the design (see https://youtu.be/rmS-oWcBZaI?).
I also mistrust the vendored packages as a way to have backdoor code in the distribution. Normal Go programs are not allowed to import code from packages cmd/vendor/..., but what if the compiler was lifting that limit for himself?
Code already decides the future of nations. My country uses voting machines which run software. People keep asking for the source code even though it would prove nothing...
I'm expecting we will finally discover a backdoor propagated since at least Go 1.4.