Exploiting null-dereferences in the Linux kernel
googleprojectzero.blogspot.com
googleprojectzero.blogspot.com
and 2 years on servers? still worth a shot. I bet it can be much faster in certain scenarios.
The article explains why it is important to not do that in general, as an oops allows debugging and recovery etc. But 2^32 of them seems suspicious.
The results about the page at address zero and the results about the page of all zeros are about equally common.
The page full of zeros is used constantly when processes allocate memory, and is set to copy-on-write.
I'm not trying to start an argument here, I think the world knows that C/C++ make it way too easy to shoot ourselves in the foot by now. I know that writing operating systems is hard and takes a long time, i've written my own prototype single and multitasking operating systems for x86_32, 68k, Z80, 6502 etc. I'm aware that Rust support has been added to recent Linux kernels, for the limited use case of writing secure device drivers. None of these things are news to me, so please don't regurgitate these points.
But given the great body of reference that is available, the enthusiasm in the Rust community for the promise of more secure operating system kernels, I'm genuinely suprised that things aren't further along. Yes I'm aware of Redox, but it seems more aimed at desktop use, and last time I tried it didn't even boot.
Projects in C/C++ seem to be making much faster progress eg. SerenityOS than the Rust community. What is holding Rust back in this area? This is a genuine question, not intending to inflame the discussion. I'm spending some time learning Rust as I can afford, but am not opinionated one way or the other yet.
Where are all the Linux replacements that I would have imagined to be up and running by now given Rust's maturity? What am I missing here? Happy to be genuinely informed.
I kind of expected there to be a bunch of projects in flight by now, ala bazaar style, with the Rust community starting to conglomerate around the strongest contenders and move them forward at a rapid pace.
Writing a kernel is like building a cathedral. To have a production ready kernel, set aside 15 years, or the equivalent in dollars, at least a couple billion. This is why stuff like Serenity, which is an outstanding achievement, is not much more than a toy.
We will have a Rust based kernel that is memory safe. Not this decade though. By that time Linux will have replaced more and more subsystem with Rust already. Remember, enterprise grade means boring and stable. It is nonsense to want a new, unproved kernel to provide that level of safety out of a language that has reached 1.0 not very long ago.
Personally, I think Linux is plenty good enough, but we have seen the best UNIX can offer. It's time to move on and try something new.
Honestly, this is a bold prediction, maybe a foregone conclusion to some, but not everybody. It completely remains to be seen.
> It is nonsense to want a new, unproved kernel to provide that level of safety out of a language that has reached 1.0 not very long ago.
Perhaps - but the Rust community has been very vocal for a very long time about how much of an improvement over C/C++ it was already before it reached 1.0. I'm not asking for an "enterprise-grade" kernel yet, just signs of prototypes starting to emerge and the community gathering round them. I realise I worded that badly in my initial post.
New C/C++/asm based hobby kernels are started (and abandoned) everyday on /r/osdev. Nowhere near as many in Rust by a wide margin. Instead the Rust community seems to have placed all their bets on an "outside in" rewrite of the Linux kernel in Rust. While I think this may well be great for Linux in the long term, its a little dissapointing that there is not more momentum for "from scratch" projects yet. Even Linux started as "as hobby, not big and professional like GNU". Which leads to my next point ...
> Personally, I think Linux is plenty good enough, but we have seen the best UNIX can offer. It's time to move on and try something new.
Don't disagree at all. That's why I'm surprised there aren't more up and coming projects in Rust with brand new designs to the degree that there are in C/C++. That's all. Maybe they are there, and I'm just not seeing them.
I suppose you might have a point: with the amount of pro-Rust posts filling almost every discussion of C and C++ over the last 5 years or more, all over the web, the time that advocates spent arguing about why C users are ignorant and stubborn could have been spent contributing to a minimal kernel.
I didn't want to say it (I have done so bluntly in the past and gotten heavily downvoted for it), but yes that is one conclusion I have reached as well. I sometimes wonder if some of the people who comment that way have ever read the LKMG.
Problems with that is that:
A) not everyone wants to dabble in writing kernels
B) there are proof of concepts see Redox
C) whatever OS you write there is almost zero chance it will make a dent. So why bother with it? It's grueling work for almost no impact.
Don't mistake a very small but very very vocal group of zealots for the Rust community at large
Yeah, but we're talking about the small group who continuously jumps into every C discussion on the web. That's still a large enough group to build a new kernel.
I think you may just be not looking where the rust people are posting. All you're really saying is that they aren't on /r/osdev, and frankly I don't find that surprising. Here are some of the more flagship projects
The biggest rust kernel/os project I'm aware of is https://www.redox-os.org/
It's a bit out of date by now, but this is an excellent guide written in rust: https://os.phil-opp.com/minimal-rust-kernel/
On the embedded side, this is a commercial project that actually "matters": https://github.com/oxidecomputer/hubris
This is a uni-kernel that I've heard about quite a few times: https://github.com/hermitcore/rusty-hermit
Naturally there's a bunch more smaller projects, and maybe larger ones that I haven't heard of.
I think one element is manpower. I don't know how many people are at good enough with Rust to write some kernel code, but I'm almost certain that it's way less than the people that can write kernel code in C or C++. C is 51 years old, C++ 38 years old, Rust 12 years old. It is possible for an enterprise-grade, production ready kernel written in Rust to exist. That still leaves the task of actually writing it.
I suspect though that Rust is also getting a lot of its fanbase from the Go crowd, or other languages from the webdev side, where there isn't really strong hardware and OSdev skills to begin with.
What would be the point? Kernels aren't swappable, and the value of an OS is in the non-kernel parts - drivers and software.
So a new kernel, even if it is enterprise-grade and production ready, won't run all the things that Linux does. Only for very niche cases, it makes sense (bare-metal hypervisors, or example).
> Projects in C/C++ seem to be making much faster progress eg. SerenityOS than the Rust community. What is holding Rust back in this area?
It's slower to write in, maybe? Unlike traditional projects where the language doesn't matter as much as the libraries available, with OS projects having a huge library doesn't help much.
For example, it might be just as fast to write a data processor in Rust as in C++, because the majority of the code is in already written libraries - you just glue together input libraries, database libraries, etc.
I doubt it. For decades C++ was mocked by Linus as unfit to write kernels in it.
But we saw resurgence with new OS like Haiku and Serenity.
You need luck to be popular, but it's not enough.
Even if not for the kernel itself, the drivers make use of C++, naturally a subset, just like the C for kernel space isn't ISO C proper.
One project that's pushing on the boundary of safety and composability is Thesus, which takes language safety to new ground by shifting traditionally OS-level responsibilities like resource management all the way down to typechecks in the language, and also explores a way of updating any core OS component on a live running system. https://github.com/theseus-os/Theseus
There's also KataOS which google just recently announced: https://opensource.googleblog.com/2022/10/announcing-kataos-...
As you note, these things take time, I agree with sibling that none of them are likely to be "enterprise-grade" or "production ready" this decade.
You're not wrong, but replace 'SerenityOS' with 'Fuchsia'. Since the latter is a much serious OS that uses Rust for its drivers.
Maybe some day Fuchsia's Zircon kernel will be rewritten in Rust and then we can all come back to this again.
UNIX-like OSes, C/C++ were just not designed for security. It is time we leave UNIX in the past and start on a clean state with the state-of-the-art and little to no legacy cruft.
I would actually be very surprised if there was anything nearly as good as Linux written in rust already. I'm not sure why a company would invest the huge amount of resources to get it done by now and, unless the language had some really unprecedented productivity, I don't think a community led protect would've had it finished by now.
I'm not sure you can take anyone (read: any troll) seriously who thinks SerenityOS is an enterprise-grade, production ready kernel.
Should we be writing new kernels for enterprise-y workloads? Why? World has kinda proven it likes things like binary compatibility. It likes for its software to work as it has, and for all the rest of the system to be predictable.
Should we be rewriting such kernels in Rust? Maybe. Where it is justified.
There are quite a few production ready Rust kernels for non-enterprise-y workloads: TockOS, Hubris, etc.
What it has is Option<T>, and I cannot turn that into a T without handling the failure case: there is literally no way to construct the code otherwise¹. One must handle the failure path. (That might be way of explicit panic/abort/oops, but it's then right there in the code: that branch will panic … and safely.)
¹(this example is using safe Rust. There's unsafe Rust too and there I can chase the null pointer all I want with that, but the parent's point is that we should be sticking to safe interfaces for stuff like this. And I'm using Rust as an example, but Option is hardly unique to Rust, heck, Rust stole the idea from its predecessors.)
(But for some userland app, aborting might be acceptable. The kernel is in a bit of a bind, since an abort — a kernel panic — means the user loses computer until they reboot, and the work along with it.)
Vs. a C pointer … all uses are more or less equally suspect; any given use, you hope the code has done it's homework for ensuring they're not NULL, and if they are, the consequence is UB. (And in Rust, and in the languages Rust steals the idea of Option from, you're only using/passing Options where "None"/null/nil is a possibility. If it's not, or you've verified or handled that at some outer stack frame, then you just pass a reference to a T, which is statically guaranteed to point to a valid object¹.)
¹again, barring buggy code using unsafe Rust, in the example of Rust, or calling into C code that fails to maintain its invariants, etc.
Take the example in the article, where the code does,
priv->mm->mmap->vm_start
while trying to generate the output for smaps_rollup. That's compilable, but buggy, C, because mmap can be null, but we failed to check for it.Vs., if mmap were an Option<T>, where T is whatever type that pointer points to. Let's say our coder attempts to write,
priv->mm->mmap->vm_start
(In some imaginary language, because C doesn't have Option, AFAIK.) The compiler would say, no, you can't "->vm_start", because "mmap" could be None (whatever you call the "nothing here" value/variant; I'm going to call it None, to distinguish it from the null pointer).In the case of unwrap, the coder could do something like (this is psuedo-code)
(priv->mm->mmap).unwrap().vm_start
It would then be obvious there is an abort there. Their reviewer would not be pleased with that, I suspect: we don't want kernel panics or oops or aborts while generating a file in /proc. And likely our imaginary coder would know this too, and when the compiler errored the first time, saying, "hey, mmap is an Option", they'd raise an eyebrow, say something like, "wait, it is? When would mmap be None?" and then proceed to properly handle that case. (E.g., by treating it as if it where the empty list.)You can't make a similar rule against null dereferences, because those happen by accident. (Unless you wrap every single pointer dereference, which is not happening.)
If you don't allow aborting, then the compiler makes you write an error-handling path that returns, and the cleanup code will not be skipped.
You're right that this is separate from the handling of the oops, which is the main exploitability that TFA is getting at, and certainly fixing one deref leading to an oops (the proc file chasing NULL) doesn't fix the other bug of "any oops can be further exploited".
But the context of this subthread is the implication that you must have some null, and some thing must happen when it is chased. That assumption is wrong, that's what the core of the comment I'm making is getting at: you can't follow a null if you don't have the possibility of them in the first place. (Or where you must have an Option<T>, you can build safe interfaces for handling that fact.)
> unwrap() would trigger the exact same bug described in the blog post with regards to reference count rollover
If we consider this instead as "an unwrap occurring during the oops handling", maybe, but it's not guaranteed that that is the case. Other aspects of Rust could similarly prevent that bug. I haven't fully grokked the latter half of the article, but I didn't think it would be necessary for the comment, as, a. the chain was about "Printing /proc/$pid/smaps is not on any conceivable performance-critical hot path." and b. followed by the question about null.
Ref-counting in Rust is often dealt with via RAII, and is safe through that, both in that RAII means the refcount is managed correctly and without input from the coder, but also Rc (and I presume Arc) will abort on overflow. I don't know if that would fully translate to kernel code, given that we might be taking refs due to the actions of userland, and that might be happening near the userland/kernel boundary and be reasonably subject to unsafe code that could very well fall prey to the same problems.
In Rust this is the ? operator: https://doc.rust-lang.org/reference/expressions/operator-exp...
The whole idea is that there should never be a way to call unwrap() if you as the caller cannot handle it gracefully. And if you do this at every step up until the UI layer, which can handle any failure as an error to be displayed in the UI, then the job is complete!
Was that supposed to be a hard question?
There is also an associated nullability sanitizer.
I use this in my own C code all the time and null pointer errors vanish if you faithfully annotate every pointer. There’s also a pragma to make non null pointers the default in a file.
GCC devs would have to be convinced to add this to GCC and then nullability annotations would need to be added to the kernel. You can then do static analysis/compile error if you do an unguarded check of a nullable pointer.
Security engineering is the field of practical mitigations - given that there are, in fact, null pointer dereferences in the kernel, mmap_min_addr and adding count limit to kernel oops provides defense in depth to help prevent them from being exploitable.
>Security engineering is the field of practical mitigations
Somehow I've managed to use bounds checking in anything I create, and I'm not even an engineer, just a hobbyist!
TLDR: However honorable the end-goal is, this blog post is not the ammo you need to push for a big rewrite of various kernel<->userland interfaces into memory safe languages.
I disagree, for profiling memory usage it's useful to get memory map data multiple times per second.