Rust's unsafe pointer types need an overhaul
gankra.github.io
gankra.github.io
I regret to inform you that x86_64 (which is probably a “major architecture”) is segmented and has three address spaces. In practice, it has two. In user code, those address spaces are normal memory and TLS and, in kernel code, they are normal memory and percpu memory.
And the proposed scheme in this article is, on first read, amazing. It would make the epic hackery in the kernel percpu code safer (in Rust, anyway). The major caveat is that, for even slightly decent codegen, the address space of an x86_64 pointer should be part of its type, not its value. Even do, this scheme still works:
ptr.with_addr(usize) -> ptr
has a type! If the input pointer is TLS or percpu, so is the output.Can you say more? I assumed the post was referring to segmentation registers, which are not an x86_64 thing.
https://stackoverflow.com/questions/6611346/how-are-the-fs-g...
but i guess the difference between "pointer to an array" and "a segment of memory" isn't actually that big?
__thread int array[100];
array[n] += 5:
Effectively computes the offset portion of array+n and then, using CPU assistance, accesses that memory relative to the segment.Making myself nostalgic for the Alpha again ... I'll be over in the corner crying into my beer.
Convenience trump's security apparently..
1. Bandaids are critical. We can't just fix the billion lines of C code out there by replacing it - we need strategies that wholesale destroy bug classes. CHERI does that. If you concretely removed temporal unsafety from all of an operating system's userland, just by recompiling, you'd have done the world a great service.
2. CHERI is very strong. It's not like some mitigation techniques, which rely on increasingly niche thread models (ASLR makes less and less sense as the world pushes towards "send code, not data" - to be fair, pax/grsec explicitly noted this 20 years ago, so it's hardly the fault of ASLR). CHERI totally wrecks spatial memory unsafety. That's fucking huge. The death of buffer overflows. Combined with other techniques like PAC, the memory unsafety threat becomes seriously less critical, or at least that's my position on it - would love to hear someone point out why I'm overly optimistic on this.
In theory you could remove all bounds checking from code and leave it up to the hardware to enforce.
It also is just sane. Like, I think containers are "sane" - they finally split the OS into disparate pieces. I think that NX is sane - why should memory be RWX by default? "A pointer for an allocation can dereference memory anywhere in the address space" is not sane, so hardware enforcing sanity is good.
Where we can enforce sanity, we should. CHERI does that imo.
Also, idk, even in memory safe languages it's not like bounds just go away. You get compile time assurances that either the bounds are gone and that's safe, or they're not gone and that's safe.
Won't any C code making use of int-to-pointer casts fail to work on CHERI? And isn't that a property of virtually every C codebase in existence? It sounds like that code will need to be rewritten either way.
> CHERI totally wrecks spatial memory unsafety. That's fucking huge. The death of buffer overflows.
You'll have to elaborate on what PAC is, but leaving use-after-free/temporal memory unsafety on the table is a pretty conspicuous hole. If the answer for that involves runtime overhead, then it will be tough to sell moving to a new architecture as opposed to just using runtime mitigations on existing architectures (or, at the limit, rewriting the important bits in Rust).
Rust doesn't have this luxury because it only defines an exact-pointer-sized integer, so it has to map usize to intptr_t and bloat up everything really badly.
https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/cheri...
And plenty of other C and C++ software as well, if you follow those links.
It's not a hole, it's just not addressed by CHERI explicitly. It's like saying that ROP is a "hole" in NX - it's just not part of the NX threat model (well, it is/was, since they knew about it - but that is not the point).
Yes, temporal safety is a concern. But removing spatial memory unsafety as an entire primitive will have consequences for practical exploitation of temporal vulnerabilities in some cases. A full exploit chain is often going to abuse both properties - though not always.
But also, CHERI plays well with other mitigations, like MTE (which PAC is a precursor, I should have said MTE).
MTE does address temporal safety. But a flaw is that it relies on pointer metadata that isn't protected against spatial unsafety. So CHERI reinforces that protection.
The point being, removing spatial unsafety in the way that CHERI does (ie: enforced by hardware) will have significant impact across the board for security.
Whether it plays out as well as I hope remains to be seen, but I think it would make practical exploitation of many temporal vulnerabilities more difficult.
> it will be tough to sell moving to a new architecture as opposed to just using runtime mitigations on existing architectures (or, at the limit, rewriting the important bits in Rust).
FWIW PAC is already deployed on every Android device afaik. I'd love to see everything rewritten in Rust but I'd still want CHERI.
> If we can find a legitimate user client that provides an implementation of getTargetAndTrapForIndex() that returns a pointer to an IOExternalTrap residing in writable memory, then all we have to do is replace trap->func with a PACIZA'd function pointer (that is, a pointer signed under APIAKey with context 0). That means only a partial PAC bypass, such as the ability to forge just PACIZA pointers, would be sufficient.
https://googleprojectzero.blogspot.com/2019/02/examining-poi...
How are you going to overwrite memory with CHERI? It cuts off a key primitive for bypassing PAC.
Yeah. So... why are you porting to CHERI if you're going to do it like that?
As a former JavaScript VM hacker, I think what I want the most is a conditional call instruction (like on arm32). The conditional flags are already computed for arithmetic, and there are already conditional jumps. JS VMs use them to check for overflow out of the small integer range (31 bits, but shifted one left, so a 32-bit arith op generates the right flags) followed by a branch to out-of-line code. That out-of-line code has to be duplicated for every different check site, since a branch doesn't keep track of where it came from. But a conditional call instruction would, either pushing the return address or putting it in a register. This allows a single, shared routine. You can then use this for all kinds of things....inline bump-pointer allocation with an automatic bailout to a slowpath, safety checks of all kinds, profiling, deoptimization, you name it.
What I wouldn't give for a conditional call instruction!
It seems like the kind of feature that competition among CPU manufacturers would drive towards an implementation out of a desire to claim faster performance on a workload that people care about.
Not sure how accurate it is, but I can imagine it being similar to parallelizing a program with independent units of works vs trying to do the same for threads that depend on each other’s output.
As a sibling commenter noted, a conditional call is basically just a macro for a (negated) conditional branch jump over a call instruction, but with a different prediction.
You can't use them for actually conditionally calling functions because functions need arguments, thrash registers, and whatever else functions do.
It never saw much use, and got removed in x64.
While only available as developer mode flag, this might mean it will be enabled by default on Android 14, as this is how such kind of features tend to be integrated into Android (1 OS generation for testing purposes and fine tuning).
https://blog.esper.io/android-13-deep-dive/#memory_tagging_e...
I find a bit ironic that for all the hate Oracle gets, Solaris SPARC has been the most successful UNIX to tame C, and has taken so long for this kind of features to spread into other platforms.
I personally think that way too often the discourse from the core team / stewards / BDFLs etc is wholly positive about their own community + the technical merits of the programming language they are representing. A healthy dose of "here are some flaws" is very refreshing and humbling, and makes me appreciate the community more.
Perhaps the original example, Hoare helped popularize null pointers, and then gave the "Null Pointers: The Billion Dollar Mistake" talk https://www.infoq.com/presentations/Null-References-The-Bill...
The creator of nodejs talking about some of its mistakes: https://www.youtube.com/watch?v=M3BM9TB-8yA
Nada Amin was part of the Scala team, and also wrote a wonderful paper about how Scala's type-system is fundamentally unsound: https://namin.seas.harvard.edu/publications/java-and-scalas-...
bradfitz was a core go team member, and wrote a post about what their net.IP type got wrong https://tailscale.com/blog/netaddr-new-ip-type-for-go/
I have no doubt there's many more examples too, but those are the ones I can think of offhand.
Semi-related, CHERI is so cool. I find it extremely promising and I truly hope that it sees widespread adoption. I really think if CHERI, or approaches similar, gain widespread adoption we will see a new era of security, similar to if not more significant than when Chrome hit the scene for desktops.
Suggestion on syntax: just use ‘->’
It has the familiarity from C already. Most rust unsafe programmers are long time C/C++ folks. You can define them to only work on raw pointers (like C) and always produce pointers (unlike C). And using it means no auto-deref magic
E.g.
ptr->field3->[0]->write(7);
Explaining it is simple:
‘.’ uses a reference and produces a reference
‘->’ uses a pointer and produces a pointer
`a.b` is of type `B`, only `&a.b` is `&B` and `&mut a.b` is `&mut B`. This is unlike your proposal of `->` where `a->b` would be `*const B`.
ptr->field in C is (*ptr).field, at which point you are in whatever C's equivalent of a "place (lvalue) expression" is. This creates a weird discontinuity where you -> for the first step and then use `.` for subsequent steps and then do the standard "just kidding, it was a pointer offset all along" thing of slapping `&` in front of it.
My proposed ~ always keeps you indirected so you just use ~ all the way until you actually want to load a value from memory (which in C would implicitly happen whenever you have a nested ->).