OpenBSD's kernel gets W^X treatment
marc.info
marc.info
Also, it sounds like it would be a massive undertaking, even for a small(ish) kernel like OpenBSD, if the kernel wasn't always written with this goal in mind. Is that the case? (I don't see a patch referenced, so I'm not able to judge for myself.)
I'm not overly familiar with all that went into this, but reading the commit logs by both Theo and Mike might help clarify some of it.
http://freshbsd.org/search?project=openbsd&q=file.name%3Aamd...
http://freshbsd.org/search?project=openbsd&q=file.name%3Aker...
http://freshbsd.org/search?project=openbsd&q=file.name%3Aamd...
http://freshbsd.org/search?project=openbsd&q=file.name%3Aker...
1. ProPolice detects most attempts to clobber the return address.
2. You can't set the return address to memory you control the contents of, such as user input, since that memory is writable and therefore not executable.
3. The remaining way to get code into executable memory is to write a file and have the program mmap(2) it. Address space layout randomisation makes finding this code difficult, even if you have the ability to smash the stack in a way that bypasses ProPolice.
The difficult cases (like the trampoline case mentioned) are usually problematic because they are programmatically writing small functions in machine code, then executing them; basically this requires the discipline to write the function and then immediately flip the page from writable to executable. Implementing a JIT compiler like the JVM would encounter similar difficulties.
More difficult, but https://en.wikipedia.org/wiki/Return-oriented_programming is a thing.
http://flyer.sis.smu.edu.sg/trustcom11.pdf
I don't know whether it's practical or not but it's definitely possible.
Such features have helped make exploits more difficult over time, but far from impossible - it all depends on the type of vulnerability, as well as things like how much interactivity exists between the attacker/the attacker's code and the target (potentially allowing em to gather data about ASLR, stack canaries, etc. before sending the final code execution bit). For example, web browsers are a very good case for the attacker, where not only is there a lot of interactivity in the form of JavaScript method calls, but a JIT usually ensures RWX pages exist; on the other side, an inetd server that spawns a new process for every request, with new ASLR offsets and stack canaries, would be pretty bad, since there is little interactivity.
When it comes to the kernel, an important attack source is userland programs (already compromised or run by a malicious user in a multiuser system) trying to abuse the system call interface. In this case, not only is there a lot of interactivity (many system calls + complex low-level device drivers, if applicable + weird CPU features + high level of control over multiple cores/threads and timing + sharing the same CPU caches etc. with the kernel), on pre-Haswell x86-64 processors, there is actually no performant way for the kernel to prevent the memory of the currently running user process from being directly accessible from it (not executable as of Ivy Bridge though), making any kind of ASLR much less useful. So while kernels can and do get pretty far by having well-written code that avoids vulnerabilities, they usually only need to give an inch for userland to take a mile. There are, however, other, less favorable attack scenarios, e.g. remote attacks on network stacks, and in any case W^X can't hurt.
http://ark.intel.com/products/27460/Intel-Pentium-4-Processo...
http://marc.info/?l=openbsd-tech&m=142122093110713&w=2
OpenBSD/i386 still has to run on systems without PAE, but it should also take advantage of the capabilities of modern processors.
As another commenter said, one of the reasons this isn't done at the outset is that some architectures (notably x86-32) don't implement W^X in hardware, so there's no pressure to be 100% clean about this. But it's rare that you have code that isn't straightforward, at least conceptually, to rework into W^X compatibility.
One of the more annoying things is fixing up relocations: if you have a call to a function in a dynamic library, and you don't know where that library is going to be loaded until it's loaded, the most obvious way to implement this is to map your program code writable and fill in addresses once you know what they are. Each place where an address needs to be filled in is called a "relocation". So when you load a dynamic library, you loop over all relocations and fill in any addresses for symbols contained in that dynamic library.
There are lots of reasons this is awful; one is that you have to go update every place in the program code. So you indirect that through a thing called the "procedure linkage table" (PLT), which contains a bunch of tiny functions that just go call your real dynamic functions, and you hard-code references to the PLT. The PLT still has to be writable, though. If you don't want that, you make a separate section of the program called the "global offset table" (GOT) that contains addresses, and you have each stub function in the PLT do an indirect function calls to a matching entry in the GOT. So the PLT is executable and doesn't need to be writable, and the GOT is writable and doesn't need to be executable.
If you want to be super paranoid, you resolve all your dynamic libraries at startup and mark the GOT as read-only ("bind now" and read-only relocations aka "relro", respectively), so that nothing is writable. But that's a different discussion from W^X.
(This telling is not very historically accurate about how the PLT and GOT came to exist, but hopefully the explanation of what they do is close enough to correct to convey the general ideas.)
W^X is a policy that a kernel can choose to implement, that if the W bit is set on a page table entry, so is the NX bit. You need the NX bit to be available in hardware for this to be useful, but hardware support for NX doesn't mean that you have to use it, let alone implement W^X.
This means that amd64 processors are backwards-compatible with kernel and userspace designs that require W|X, even in long (64-bit) mode.
once:
mov byte [once], 195
; ...do something here...
ret
I have used this technique in applications-level code, where it is significantly more efficient (both smaller and faster) than the alternatives when this "once" function will be called many times. I think it is always important to remember that while W^X and other restrictions have security benefits, they also have downsides in limiting some interesting creativity and the potential to exploit the full abilities of the machine.However, this does kinda ignore one of the big focuses of the OpenBSD project. They tend to shy away from such clever hacks in the name of readability and auditability. While it's definitely a neat way of ensuring your code is only executing once, it becomes a hassle when you have to port it to other platforms. Keep in mind that OpenBSD ports to as many platforms as possible because the subtle quirks of various platforms will often tickle out rare bugs to become more repeatable. In this case, your replacing the once function with a return is dependent on x86, so wouldn't work on the many other platforms that OpenBSD runs in.
Perhaps it would help to see them "in action" here is a link to the linux emulation layer in freebsd, the MD chapter four specifically describes i386 (this is how you put syscall parameters on, and off, the stack on a i386) and the MI chapter five is a pile of structs that would be used by any emulation layer (NPTL, TLS, the joy of futex'es's (a linux thing that is kind of a mutex cache for speed, sorta kinda), and good luck with the ioctls).
https://www.freebsd.org/doc/en/articles/linux-emulation/inde...
http://freshbsd.org/search?project=openbsd&q=SMEP&committer=...
http://freshbsd.org/search?project=openbsd&q=SMAP&ommitter=j...
Otherwise it's a seriously impressive feat.