HNHacker News
TopNewBestAskShowJobs

jrtc27

173 karma · joined February 10, 2021

submissionscomments
jrtc27··on Arm releases experimental CHERI-enabled Morello board
CHERI is orthogonal to virtual memory, and the two complement each other. You still want virtual memory so you can do the usual paging tricks, copy-on-write, sharing of read-only pages, and so on. Plus the fact that there is a single page table entry for an address that affects all accesses is crucial for our experimental temporal memory safety implementation (see https://msrc-blog.microsoft.com/2022/01/20/an_armful_of_cher...). There's nothing stopping you from using segments instead of flat address spaces with page tables, but it's not really related to CHERI, you still have the same trade-offs as you do on conventional architectures.
jrtc27··on Arm releases experimental CHERI-enabled Morello board
Similarly if you really want mass market adoption then you need a Windows port, otherwise most consumer PCs will remain without it, and for mobile adoption you want an iOS port (though Android does at least contribute a sizeable chunk). Porting FreeBSD does, however, not just serve as a PoC but also let you port all the standard third-party software that runs on all Unix-like OSes (most ports need few if any changes, but with tens of thousands of software packages out there it does add up if you want a full set of packages available), as well as being a reference implementation for other CHERI OSes to use when being ported since we'll likely have already encountered most of the friction points they do. Plus FreeBSD has its Linuxulator which provides a binary compatibility layer for Linux binaries, so you could even develop parts of a CHERI GNU/Linux userspace on top of that without a real CHERI Linux kernel implementation (we have a proof of concept port of the Linuxulator, but it's not currently fully fleshed out, in part because there wasn't even a proper CHERI Linux ABI defined by Arm at the time).
jrtc27··on Arm releases experimental CHERI-enabled Morello board
The C startup code (for statically-linked binaries) and run-time linker (for dynamically-linked binaries) carve up initial capabilities provided by the kernel into capabilities that cover the various global variables and function pointers needed by the program and libraries, similar to how pointers are initialised for position-independent code (more complex, but same principle, just scan through all the relocations and apply them). When you mmap(2) memory from the OS, you get back a capability with bounds covering that memory. When you malloc(3) memory from your libc, it finds space in an existing mapping, takes that capability and restricts its bounds to the allocation size. When you take a pointer to a stack-allocated variable, the compiler inserts an instruction to set the bounds of that capability to just the memory it allocated for that variable. Every pointer, whether "language-level" (what is exposed in the language) or "sub-language-level" (the pointers in the implementation, like return addresses on the stack or the stack pointer itself), is a capability, and all you need to do is insert a bounds-setting instruction at the point of allocation to restrict its bounds. So your libc's malloc needs modifying, as does your kernel, but your C program that calls them just needs to be recompiled for the pure-capability ABI.

Edit: To answer the first question, yes, that is the primitive which enables CHERI to be used for in-address-space compartmentalisation rather than relying on an MMU for process-based separation and all the overheads that come from context switching address spaces.

jrtc27··on Arm releases experimental CHERI-enabled Morello board
The simple answer is that it actually works for real-world software, is microarchitecturally feasible and flexible, and architecturally enforces non-forgeability (which is crucial allowing in-address-space compartmentalisation of distrusting software). Most schemes that take the metadata-table-on-the-side approach fall down on those last two points. MPX is particularly notorious for tanking performance, having race conditions (because loading the bounds is not atomic with loading the address) and having an extremely limited number of bounds registers (I think 4? which is even worse than the highly constrained register set of 32-bit x86) so you're constantly spilling/reloading bounds data from memory. I don't think any of them have been shown to work across the entire software stack from the kernel to core userspace runtime parts to graphical desktops like KDE.
jrtc27··on Arm releases experimental CHERI-enabled Morello board
Not really, because it gets traded off with the increased memory pressure due to the larger pointer size, and it'd likely be workload dependent. It's not something we've explored to date beyond hypothesising that it could be a good thing.
jrtc27··on Arm releases experimental CHERI-enabled Morello board
I don't see why it implies that. "Arm releases experimental DDR5-enabled $NAME board" wouldn't make it sound like DDR5 is only for Arm, so why would "Arm releases experimental CHERI-enabled Morello board"?
jrtc27··on Arm releases experimental CHERI-enabled Morello board
Windows is likely a big task for the same reasons as SMAP (https://github.com/microsoft/MSRC-Security-Research/blob/mas...). XNU should be comparable to FreeBSD, which CheriBSD is a fork of, as both use Mach's VM for memory management and have a bunch of shared code in various places, but userspace is more of an unknown quite how much effort it'd be (you'll need to port Objective-C and, now, Swift, for example). For Chromium we have ported WebKit, so I'd imagine Blink isn't too dissimilar. V8 is likely interesting, though we have a version of WebKit's JSC JIT for Morello, which gives confidence in V8 being doable.
jrtc27··on Arm releases experimental CHERI-enabled Morello board
Software doesn't need to "adopt the instructions", it just needs to be recompiled in the same way as you compile it for a new architecture (CHERI is effectively like the 32-to-64-bit transition in that sense). Yes, having capabilities allows you to bring memory protection to the MMU-less embedded space (see for example the now somewhat old paper https://www.cl.cam.ac.uk/research/security/ctsrd/pdfs/201810...).

Yes, if you attempt to access outside the bounds of a capability you will deterministically crash. This is true even if you do have virtual memory and there is memory there.

Yes, the use of CHERI to protect unsafe code in memory-safe languages like Rust is of interest to us. There is also the possibility of being able to remove some of the compiler-generated bounds checks by using the capability bounds instead, though some care is needed to preserve the precise semantics (but some may also be happy to slightly change the semantics if it means they can all be removed and potentially improve performance).

jrtc27··on Arm releases experimental CHERI-enabled Morello board
We do have formal proofs of various security properties at the architectural level that consider the entire architecture with all its complexities and warts. Speculative execution is of course a concern (and is an active area of research for us), though our belief is that the bounds information now present at the hardware level allows it to be tamed. Another concern is the interaction between undefined behaviour and CHERI; the former needs to be sufficiently constrained in order to not inadvertently turn code that would be memory safe with a naive CHERI compiler into code that is not. We also know there are still memory safety-like issues we can't protect against; we can stop pointer injection, but we can't stop tricking programs into copying the "wrong" pointer somewhere, or type confusion bugs that result in using pointers in an unintended way. Many of those exploit chains today rely on exploiting some other memory safety vulnerability we do protect against, but we cannot predict if people will come up with alternative approaches that avoid those in a world with CHERI.
jrtc27··on Arm releases experimental CHERI-enabled Morello board
We also have a CHERI-RISC-V specification (https://www.cl.cam.ac.uk/techreports/UCAM-CL-TR-951.pdf), with support in CHERI LLVM, CHERI QEMU and CheriBSD, plus three open-source FPGA implementations (https://github.com/CTSRD-CHERI/Piccolo, https://github.com/CTSRD-CHERI/Flute, https://github.com/CTSRD-CHERI/Toooba) that span various parts of the microarchitecture design space, and it is the platform we use for our own research on architecture and microarchitecture. But for various reasons (e.g. proximity to the university, existence of competitive microarchitectures several years ago, ISA and ecosystem maturity, enthusiasm and interest on their part) Arm was the right partner for this program.
jrtc27··on Arm releases experimental CHERI-enabled Morello board
Our work is based on FreeBSD as its tight integration makes it much easier to manage forking in a research setting, compared with the umpteen different repositories you need to fork and keep in sync to build a Linux distribution. Arm have a minimal Android stack and are working on a Linux distribution (but their current Linux kernel implementation does not enforce capability protection, it's done by a userspace wrapper, and only a select number of binaries in the Android image are pure-capability, many are still plain AArch64), initially based on musl, but it's still a long way behind where we are on FreeBSD where we have (almost) all of the userspace and kernel ported as pure-capability code (the "almost" is because we have not yet invested the engineering effort in porting DTrace and ZFS, but both are on our roadmap as they're important for real use).
jrtc27··on Arm releases experimental CHERI-enabled Morello board
Nobody's claiming it's "hack-proof", that would be foolish, just that it removes certain classes of vulnerabilities that are the majority of CVEs for code written in memory-unsafe languages, thereby reducing the attack surface. Independent analysis by both Microsoft and Google has shown that's around 70% of vulnerabilities, which still leaves around 30%, but is a big step forward.
jrtc27··on ARM ships ground-breaking Morello secure processor board
"ARM has developed a prototype architecture based on the Cortex-A core" isn't quite right. Armv8-A (well, Armv8.2-A in this case) is the architecture, Cortex-A is a family of implementations of that architecture. And the Morello implementation (called Rainier) isn't even directly based on a Cortex-A core, it's based on the Neoverse N1, which is itself then derived from the Cortex A76.
jrtc27··on You can't copy code with memcpy
You can allocate memory in another process on Unix too: use ptrace to make the other process call malloc (use PTRACE_SETREGS to set PC to malloc and the first argument register to the number of bytes, then intercept the return).

GDB will use this if you tell it something like `p foo("bar")`, as it needs to allocate memory for that string somewhere.

jrtc27··on Children's Risk of Serious Illness from Covid-19 Is as Low as It Is for the Flu
The graph is deceptive; it's a semi-log plot, so the linearity is really exponential.
jrtc27··on WebAssembly and Back Again: Fine-Grained Sandboxing in Firefox 95
Most C and C++ code is perfectly happy with fat pointers provided you take care in how you implement them, especially around (u)intptr_t. For CHERI we see a very tiny % of LoC that need changing; e.g. http://www.capabilitieslimited.co.uk/pdfs/20210917-capltd-ch... documents a recent case study in porting a KDE desktop stack (X11 itself, Qt, other standard graphical desktop libraries, Plasma, Dolphin, Okular) and across the 6 million LoC only 0.026% needed changing.

So, if you pick your fat pointer implementation poorly, then yes, you run up against both the standard and de-facto C, but if you take care then the vast majority of C code just works, and we're (uniquely?) positioned to be able to prove that by having a FreeBSD-based kernel and userspace, and graphical desktop stack (plus an adaptation of WebKit's JSC JIT) all built with a compiler that maps C pointers to capabilities that enforce fine-grained bounds.

jrtc27··on Rust: “Move fast and break things” as a moral imperative
CHERI (http://cheri-cpu.org) provides a spatially-safe C and C++, and also heap temporal safety (specifically it prevents use-after-reallocation, which is the actual vulnerability, since use-after-free doesn't matter if freed memory is never reallocated). Most code requires few, if any, changes (0.17% LoC in our fork of FreeBSD, which includes the kernel itself and all the low-level runtime libraries), with the changes tending to be due to people conflating pointers and integers (i.e. use uintptr_t not unsigned long/uint64_t for storing a union of a pointer and an integer, or a real union, and use size_t not uintptr_t when you mean a plain integer, though the latter can sometimes still work, just a little less efficiently and likely with some compiler warnings). It doesn't solve the concurrency issues, but unlike CHERI C/C++ those require invasive changes to the language and thus code.

We have a technical report that gives an overview of CHERI and describes how to write good C/C++ that doesn't use dodgy idioms that break in CHERI C/C++ if you're interested at https://www.cl.cam.ac.uk/techreports/UCAM-CL-TR-947.html. Arm are also working on a prototype of Armv8-A with CHERI, dubbed Morello: https://www.morello-project.org.

← PreviousPage 2 of 2