(You can also build with clang, but only because clang deliberately aims to support most gcc extensions.)
is the benefits of "modern C" worth compile times 2-3x times longer?
Are you able to give hints that would guide us in thought exercises?
You have to be very large to notice it, but the likes of Facebook, amazon, and Google have massive warehouses almost entirely filled with computers. It doesn't take much to see how their power bill can add up.
The Distribution collectors, they brag about speed, and optimization is very important to them. They want to be able to test changes very quickly.
"Jun 30, 1987 — The Cray is part of a $20m installation used by Apple's Advanced Technology group and consists of four CPUs operating at 9.5nS per cycle,"
Now we have distcc, that can both massively parallelize compilation, and optimization.
Googles team probably shows some slight improvement, just to justify their work, but in the long run, its just as sloppy as industry wide coding is.
Microsoft is the absolute worst, along with Apple.
I would guess that you're referring to either "How ISO C became unusable for operating systems development" ([0]) or "How One Word Broke C" ([1]).
[0]: https://arxiv.org/abs/2201.07845 , most recent HN discussion at https://news.ycombinator.com/item?id=30022022
[1]: https://web.archive.org/web/20210307213745/https://news.quel... , HN discussion at https://news.ycombinator.com/item?id=22589657
So “old C” in such a case would need to mean “an old C implementation” (or possibly a new one, but simple or configured to behave like an old one), something like GCC 2.8 maybe, and nobody’s using that on desktop. So the language standard version should be mostly immaterial, and it’s not like the C89-to-C17 difference is anything like the yawning C++98-to-C++20 chasm. (This is a carefully phrased statement: C99 had complex numbers, which are annoying, and variable-length arrays, which are a significant change, but C11 demoted both to optional features.)
I mean, you get more security by default and security is pretty darn important for kernels...
For example, Rust is frequently said to have long compile times. There are a number of interesting reasons why that is often true, but if we ignore the pathological cases then what we find is that borrow checking takes a significant fraction of that compile time. This is a big trade–off between features and complexity that the Rust language made very deliberately: the advantages of the borrowing rules are what makes Rust such a great language. The cpu time spent checking that those rules have been followed are a small price to pay, but not a negligible one.
The ladder's being moved a millimeter.
As compiler technologies evolve, it becomes not only a necessary evil but rather a better course of action to trust compilers rather than humans tiptoeing around security risks masquerading as language idiosyncrasies.
Erm. I was always under the impression that he had actually done it, and that his presentation was a historical anecdote. Is that not the case?
It's more work - you have to figure out a sneaky bug and write a legit looking patch - but anyone can do it. You don't have to be in a position of power already (e.g. being the Debian GCC packager) so overall it is much easier.
For example, I run Gentoo Linux, and haven't reinstalled my OS since 2004 or so. That means that, modulo a few binary packages, I have a direct source lineage to the state of Linux in 2004. If you want to pull off that attack against my system (and you didn't already back in 2004), you'd have to tamper with source archives. That would both imply changes that are easy to analyze (more than binary patches), and it would involve changing the archive hashes in the Portage tree. That tree is in Git, which means that it would create an immutable public record of what happened (Git is the original blockchain, remember), modulo forced pushes which people would, again, notice all over the place.
In practice, if you want to persistently backdoor a new system (supply chain attack), it's usually easier to do that in hardware or firmware than trying to do a RoTT attack on the distro and its compiler. In fact, it's users of binary distributions (or proprietary OSes) that should be more worried, as it is much easier to do a binary-based RoTT attack that self-updates to handle new versions consistently when all your users run the exact same binaries. Source code users should be more worried about compromise upstream than local persistence. And those attacks are a review / auditing issue, unrelated to RoTT.
In the end, if you are worried about being personally targeted, it's easy enough to make that impractical by re-bootstrapping your computing from an unpredictable source (e.g. walk into a random shop and buy a PC, walk into a net cafe and download your favorite distro and check the hashes there). And if you are worried about large-scale attacks, RoTT style ones aren't practical without someone somewhere noticing; you should be worried about traditional compromise instead.
Thompson's hack relies on the Halting Problem, and the space for deviousness within the Halting Problem is infinitely large.
C11 added insecure Unicode identifiers, and you won't find them in reviewing them manually via email. You need a special linter to detect such Unicode attacks. Or a proper development environment.
Compiler development usually evolves into more bugs, not less. Just now they caught up with the hundreds of bugs they added with gcc-9. Not talking about const, restrict, and strict aliasing, and the still not existing -Oboring for the kernel. Would you dare to use -O3 and -flto and -fstrict-aliasing in the kernel?
Heck, Linux already uses GCC plug-ins to implement some fancier security stuff; it is silly to think they can't handle forbidding Unicode identifiers.
Actually, as a compiler writer, I'd go a little bit further and point out that Linux itself isn't even written to the gnu C89 very well; it's often written to a "C is portable assembly" view of the language, which results in nasty grams and invective being hurled at compiler writers if they compile the C specification correctly and not according to the "proper" assembly the code author thought they were getting.
One of the benefits of more modern language revisions is that they actually tighten the wording on a lot of the more ambiguous parts of the specification--C11 in particular adds a much more comprehensive memory model that's very shrug in the older revisions of C.
Their hesitancy with newer standards is understandable, when viewed against that backdrop
More to the point, though, the only changes to undefined behavior in the C specification in newer versions (compared to C89) are either clarifying things that were already undefined behavior (e.g., INT_MIN % -1) or actually making some undefined behavior well-defined (e.g., allowing type punning via unions).
[1] http://www.open-std.org/jtc1/sc22/wg14/www/docs/n2464.pdf
Older standards said "it is implementation-defined whether the old object is deallocated". This didn't really work well:
- if realloc(..., 0) returns NULL if it freed the object, then you have confusion with error cases. strtol already has this kind of interface and it's unusable
- if realloc(NULL, 0) returns NULL and does nothing it does something different than malloc(0). Some chose to make it do something different, some chose consistency with malloc.
- if you choose consistency with malloc then realloc(NULL, 0) likely will end end up inconsistent with realloc(ptr, 0) where ptr is not NULL. On BSDs the two are consistent but also different from any other platform, so portable code could not rely on realloc(..., 0) doing something known: either you had possible double-free bugs on some platforms, or you had a memory leak.
With that said, we still have options even when moving to a newer standard. Many new language features can be machine-translated to C89 if needed (similar to how we have protoize / unprotoize, though not always as seamless). If we need to keep a bridge to C89 the kernel authors could hold to a subset of new language features that are amenable to machine translation.
Why does the complexity of the OS, or even its ability to be compiled by multiple compilers, matter for "trusting trust" attacks?
As I've always seen it, the problems and solutions all exist at the compiler level. The whole premise relies on starting with a binary compiler that you are expected to trust. The source is also assumed to be safe and un-tampered for purposes of this discussion because that's an entirely different issue.
The solution, of course, is to build the compiler itself with different compilers. If you build the same compiler with two or more different compilers, then use that to compile itself, you should be able to with the right options (see the work that has been done on reproducible builds) get binaries that are equal or close enough to easily compare any differences.
At that point if they are the same then you know either it's good or both of the upstream compilers were also compromised.
If for whatever reason that's not practical at the top level compiler the same concepts apply going further back in history until you get to some early compiler a bored grad student wrote in the '80s in pure ASM.
As a result from a practical sense I don't really see "trusting trust" attacks to be that big of a concern. It's always possible to work your way back down the tree of software until you get to a point where you can actually trust a compiler and then build your way back forward from there.
If a compiler starts depending on its own tricks and gets to a point where it can only be successfully compiled by itself, then there are reasons to be suspicious. Even then you'd just have to have the last version to be able to be built by other packages as one more stop along the way to trust.
I was always under the impression that the "message" of Reflections on trusting trust was that you have to trust someone at some point.
Man I so vibe with that. But the new Cs do have a ton of new features and, more to the point, I trust Linus to make this kind of decision more than just about anyone.
https://www.win.tue.nl/~aeb/linux/hh/thompson/trust.html
The paper's a pretty entertaining read:
> First we compile the modified source with the normal C compiler to produce a bugged binary. We install this binary as the official C. We can now remove the bugs from the source of the compiler and the new binary will reinsert the bugs whenever it is compiled. Of course, the login command will remain bugged with no trace in source anywhere.
[0]: https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_Ref...
Ken Thompson, Reflections on Trusting Trust.
EDIT: I don't think I've seen 5 nearly simultaneous replies sharing the same link before.
LOL I was searching for a non-PDF link and delayed the reply. It would have been 6 simultaneous answer.
How do you prove your stack is secure?
[1]: https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_Ref...