Strange Loops: Ken Thompson and the Self-referencing C Compiler
scienceblogs.com
scienceblogs.com
http://cm.bell-labs.com/who/ken/trust.html
Sort of disappointing that the blog entry doesn't bother linking to the original, which is at least as well written.
"I have watched kids testifying before Congress. It is clear that they are completely unaware of the seriousness of their acts. There is obviously a cultural gap. The act of breaking into a computer system has to have the same social stigma as breaking into a neighbor's house. It should not matter that the neighbor's door is unlocked. The press must learn that misguided use of a computer is no more amazing than drunk driving of an automobile."
Nowadays, I here more about the opposite problem -- using overly broad, non-technical legalese to convict kids of "hacks" that aren't really hacks. I wonder what Ken would think of Snowden, Barrett Brown, weev et al?
My favourite part was "This is a deep concept. It is as close to a 'learning' program as I have seen" since getting a compiler to recognize when, where, and how to inject (self-propagating) backdoor code seemed practically impossible.
It also expressed the idea of "self-referencing" more succinctly than the backdoor trick.
The cycle can be broken by using a different compiler with a different pedigree in the bootstrap process. In fact, this suggests a way to detect the back door:
Given Compiler A which is the previous generation, and Compiler B which has a different pedigree, we want to generate Compiler C and detect if it is compromised:
A compiles C which compiles C1
B compiles C which compiles C2
C1 and C2 should match. If it does, and A is "known good", then C is now "known good".
Of course, this can be compromised, too, but it will make it very much harder to do so.
Remember that C1 and C2 are generated by two different compilers generated from the same source code. The same source code should generate the same binary.
I.e. take the iteration one more time.
(I actually do this as part of the test suite for the Digital Mars C++ compiler - it takes two iterations until the binaries match exactly.)
(Plus a backdoor could easily be set to run only when optimizations are on. Almost all production builds that one would want to exploit should be built optimized, and it would make the exploit a bit harder to spot in compiled binaries.)
A compiles C and produces C1. C1 compiles C and produces D1.
B compiles C and produces C2. C2 compiles C and produces D2.
D1 and D2 should be identical.
I grant that this is not a very practical attack, but it can be made arbitrarily hard to detect, based on the threat model the original attacker (who compromised the C compiler) uses.
So you run your compiler, and it punches cards for you. You then turn off the machine, remove the cards, and verify them. If that checks out, then you boot the machine with the cards again. Anything that persists the reboot has to be on those cards, and is therefore subject to uncompromised inspection.
Not really. The more binaries you have to be able to subvert for it to work, the harder the attack becomes. There are lots of binary diff tools out there, lots of compilers, lots of assemblers, etc. Writing an attack that could detect them all and defeat them all would be so hard as to be impractical; being able to generically recognize something like "this is a compiler" or "this is an implementation of AES" would probably involve solving the halting problem. And if you can't solve it generically, you have to start adding bigger and bigger suspicious binary blobs in which encode all of the patterns for all of the programs you're looking to subvert.
Remember, the NSA does not have infinite resources, nor can they solve unsolveable problems. There is a practical limit to how paranoid you need to be.
Practically, subverting general purpose computation, like subverting a compiler or subverting general purpose instructions on a CPU, is likely to be too easy to detect and too hard to implement to be worth it.
Much easier is subverting a random number generator. That's incredibly hard to detect, and easy to implement. AES encrypting an incrementing counter with a key known only to the NSA looks an awful lot like a random stream, but with that key they can trivially figure out where it is in the stream and what will come next.
That's why Intel's insistence on getting the Linux kernel to blindly trust the RdRand instruction was quite worrying[1], while applying only a bit of paranoia and using David A Wheeler's approach to defeat the trusting trust attack[2] is likely sufficient to have faith in things like your compiler.
1: https://en.wikipedia.org/wiki/RdRand 2: http://www.dwheeler.com/trusting-trust/
But that is not the point. The point is that against a determined adversary, you can not trust an arbitrary program, even if that program is compiled from source, unless you can also trust every other component on your system.
The point then, is to encourage that "bit of paranoia" and make people understand that thinking you're safe just because you can read through the source or have another mechanism for producing or obtaining what you might think is a pristine, safe copy of a piece of software is a false sense of security.
While this particular attack to my knowledge was only a thought experiment and I've never heard of it occurring in the wild, note that at least a few viruses for example propagate by modifying binaries to act as carriers, and often modify the system to obscure their presence (e.g. report incorrect file sizes, and not read back the modified data). This is largely the same threat, and in many ways far more practical because it doesn't require people to take the step of trying to recompile applications.
Having the source (or a "known good" source of your application) does not help you, as your binary is infected when it is handled by the compromised system. Taking checksums etc. on the compromised systems may not help you.
I wonder how well that would work in practice. There are actually quite a few bugs in C compilers[1], and compilers are complicated enough that I could see them hitting some of them when they are compiled. Has anyone done this before?
1: http://www.stanford.edu/class/cs343/resources/finding-bugs-c...
This is utterly false and reveals a deep misunderstanding of compiler technology. Aside from optimizations and exactly how they are implemented in each compiler there are still many more or less arbitrary choices that a compiler needs to make in order to translate source code into machine code. Address layout being one of the most prevalent. Indeed, these choices are so arbitrary and lead to such a high degree of difference in the binary forms that the same compiler running on the same code will often generate different binary output. And the same compilers running on very slightly different code bases will quickly generate substantially divergent output. This is why programs like bsdiff and courgette exist, because comparing binaries is actually an enormously non-trivial problem.
I've written many professional compilers, front to back, and I use the binary difference technique to verify that the compiler is capable of exactly reproducing itself.
If you've got a compiler that generates different binaries depending on the time of day, the address the compiler was loaded at, or something else that is not the compiler switches + source code provided to it, you've got a compiler with serious QA issues.
More importantly the salient point was about comparing the binary output of different compilers.
While it's certainly possible to create tools which make it make it possible to determine if the binary output of different compilers are effectively the same such tools are very non-trivial to create. The idea that different compilers are likely to produce exactly identical output is sheer fantasy.
No, it wasn't. I explained it (apparently badly) 3 times now. There's another iteration of bootstrap compiling in there before the output is compared.
But yes, i think this method does work. You'd have to trust all the pipeline programs used in between.
If I was trying to detect a "trusting trust" type backdoor, one thing I might do is model the control flow of the executable as a graph, and use something like approximate graph isomorphism to detect additional large blocks of code.
The only way to tell if a compiler will put a backdoor into certain programs is to read the disassembled source.
So: the difference you'll find isn't the login-backdoor itself, it's the code that waits for a login program (or another compiler) to be compiled, and inserts that backdoor. This avoids the problem of having to know what the login program or whatever is.
'Known-good' is exactly the problem Thompson's essay is describing. There is no 'known-good'. Instead, you have decided to root your chain of trust with a compiler you call 'known-good' (say, B). But ultimately, that trust is arbitrary: your compiler B relied upon un-investigated components at some point in its heritage: a hex editor, a disk drive controller, a CPU, an LCD screen, etc.
The best you can do is reduce the trusted base to the smallest amount, then make explicit your trust relationships as much as possible. Ideally, the trusted base would be physics and logic with a metaphysical certainty of the universality of those laws; we're a long way from that situation :)
So yes, you're still relying on things below that. That doesn't make the exercise 'arbitrary' or pointless. Good risk management is reducing the chances of the most probable attacks, and a compromised compiler is (ISTM) a much more likely attack than e.g. a compromised CPU.
Seems to me that putting theoretical perfection too far ahead of pragmatism just gives the (wrong and damaging) impression that not being able to solve all trusting-trust issues entirely means it's not worth their time trying to solve the more tractable ones.
However, Thompson's essay is both a practical attack and an abstract idea; I find the notion of trust with self referential systems more interesting than the specifics of how to detect a backdoored compiler.
With this assumption, to write a compiler that's safe, you can't run it through an existing possibly-compromised compiler. You would have to bootstrap your safe compiler, as if for a totally new language.
So you'd have to initially write it in a different language. You probably can't write it in most existing languages, e.g. Python and Java are both implemented in C. Because, if you tried to write a safe compiler in Python, a sufficiently smart [1] C compiler would have been able to tell that you were compiling a Python interpreter when it compiled /usr/bin/python, and inserted a backdoor into that Python interpreter which will trigger when the interpreter is interpreting a C compiler written in Python.
You'd basically have to consider any code that has ever passed through an automated tool to be potentially backdoored, so you'd have to start writing in machine code (no assembler allowed, of course, because it's probably written in C). Of course, you could program an assembler in pure machine code (or using a potentially-tainted assembler, verifying its output by hand).
[1] It'd either be a general artificial intelligence of human level programming ability, or some kind of magical oracle.
Note, I can't find the actual article now - I think it may have been in one of the early issues of Phrack[2].
[1] http://www.intel-assembler.it/portale/5/Write-an-assembly-pr...
Let your imagination run if someone had done this at some point to a build of GCC being used by Linus...
I don't think any buildchains currently implement anything like this though.
If it were a GPL-licensed compiler (e.g. GCC), they wouldn't need to distribute the compiler-code changes, since they would only be using it internally and not distributing the binary for the compiler itself.
Of course, they could just modify and use a shared-library of the GPL'd code or whatever. At least, that's my understanding, could be wrong...
My understanding is, under the GPL, if you distribute the binary, you need to distribute the source.
The “source code” for a work means the preferred form
of the work for making modifications to it.
...
The “Corresponding Source” for a work in object code form
means all the source code needed to generate, install, and
(for an executable work) run the object code and to modify
the work, including scripts to control those activities.