How to Build an Evil Compiler
awelm.com
awelm.com
This is mistaken, as there exists a second Rust "compiler" known as mrustc (written in C++). I put "compiler" in quotes, as it doesn't implement all of Rust's static analysis, so it's more like an alternative code generator. That said, it still suffices to build the compiler itself (it was originally devised as a rustc bootstrapping tool), and it has been used in the past to successfully perform DDC.
Look, as this article demonstrates, it's not hard to build a backdooring compiler. Even if you want to build it in a more robust way than checking filenames, it's really not difficult: it's (admittedly complex) pattern matching, and quite a lot of optimization in fact boils down to pattern matching. The problem is that the pattern matching you'd need to do to get the everything-is-backdoored scary effect is brittle as fuck.
Compiler output tends to be effectively nondeterministic. I mean, the goal of the compiler is to produce completely deterministic input, but very subtle changes can have cascading consequences. (I say this as I am trying to fix a test for LLVM's opaque pointer changes). Even something so simple as figuring out how to make bit-equivalent reproducible builds with the reproducible builds initiative took a few years to really get going, since there are so many things that are effectively random that you wouldn't consider at first (e.g., iterate over all files in a directory).
(My examples on x86 involved changing JGE to JG, or JL to JLE, corresponding to changing >= to >, and < to <=, in loop conditions.)
Combining this with the trusting trust attack, you could have a self-perpetuating bug in the compiler plus a bugdoor in other software. The pattern match for the other software does not necessarily have to be super-specific in that case.
I would definitely agree that this wouldn't survive that many generations of software evolution without active intervention. It definitely wouldn't survive a change of programming language or target machine architecture, for example.
The Ken Thompson hack is not undefeatable. You can detect it using a cross compilation technique comparing the binary output with a clean complier. I think you have to about 4 compilations to figure out if you're infected, but then you don't know which one is infected and which one isn't. You will need more data points to compare. Disassembling the binary would help as well if you know what you are looking for.
The current best known defense is Diverse Double-Compiling (DDC), introduced by David Wheeler in 2009. To briefly summarize DDC uses different compilers of the same language to test the integrity of a chosen compiler. In order to pass this test the attacker must have modified all the selected compilers beforehand to insert backdoors into each other, which is a decent amount of work. DDC is a good idea but it has 2 shortcomings that come to mind. The first is that DDC requires all selected compilers to have reproducible builds, meaning that each compiler always generates the exact same executable given the same source code. Reproducible builds aren’t very common because compilers by default include things like timestamps and unique IDs in their builds. The second shortcoming is that DDC becomes less effective for languages that only have a few compilers. Also DDC can’t even be applied to newer languages like Rust with only one compiler. In summary, DDC isn’t a silver bullet and the Thompson attack is still considered to be an open problem.
"David A. Wheeler’s Page on Fully Countering Trusting Trust through Diverse Double-Compiling (DDC) - Countering Trojan Horse attacks on Compilers"
There are malware analysts who are good at finding sophisticated malware in binaries. They could probably locate suspicious code that could be obfuscated malicious code.
Other than that, if you can't trust any of the compiler vendors (which makes things like checksums on an HTTPS website useless), you'll have to write your own. What about the firmware of the machine you're writing your clean code on? Taking this argument to its logical conclusion makes the idea of developing even simple software astronomically expensive.
Nobody can realistically know if all of the components from hardware to software is clean. And the more technical you are, the more you realize how many attack surfaces there are. It's impossible to verify everything and you just have to blindly trust that it's all safe.
You assert that the other compiler won't contain the exact same backdoor, not that it contains no backdoor.
> different compilers could produce different, equally valid instructions, such as debug vs. release builds
The binaries you're comparing are output by two instances of the same compiler codebase (each instance created by a different compiler). So as long as that compiler is deterministic, each run should have the exact same output.
This is my primary issue with closed source software, binary blobs, etc; trust.
I mean some sort of "evil CPU": even if all the software is clean, it will act differently on some inputs (read: some pattern of instructions) to produce output in a way that it's indistinguishable at any user-level "views" without a logic analyzer analyzing all the bits in and out the CPU in physical level as electrical signals which is beyond reach practically, while still performing some evil tasks. A similar idea can be expanded to an "evil RAM" or "evil bus" but you probably get it.
I'm pretty sure there are already backdoors already in many CPUs that we aren't aware of, but I was wondering if this particular type of attack has ever been spotted in the wild.
Yeah, how is this going to generalise to _any_ program that gets compiled without introducing bugs in the executable? If you can write a backdoor that can accept _any_ program and compormise its integrity you can solve the haulting problem, can't you?
So this only works if the bad compiler has access to the source code that its compiling before hand and a _hand crafted_ hack is made for it. With opensource projects like the Linux kernal i can see this is a problem but for everyone else, meh.
Also OSS makes up most of the modern stack, so access to source code is a given. And hand-crafting a backdoor when you have the source code is trivial because you can literally change anything you want with confidence.
This is something that could easily occur with scripting languages, backend systems, open source, closed source, etc.
Basically any black-box system that takes in some input could pre-manipulate the input yielding an unknown/unexpected output.
IMO thats why it's a scary attack. It's a really simple idea and there are so many ways to apply it
https://www.quora.com/What-is-a-coders-worst-nightmare/answe...