QEMU VM Escape
blog.bi0s.in
blog.bi0s.in
1. Production VMs almost exclusively use tap networking, not slirp. This CVE mostly affects users running QEMU manually for development and test VMs.
2. Slirp (https://gitlab.freedesktop.org/slirp/libslirp) is part of the QEMU userspace process, which runs unprivileged and confined by SELinux when launched via libvirt. To be clear: this is not a host ring-0 exploit!
3. Getting root on the host or accessing other VMs requires further exploits to elevate privileges of the QEMU process and escape SELinux confinement.
More info on QEMU's security architecture: https://qemu.weilnetz.de/doc/qemu-doc.html#Security
For a more detailed overview of how QEMU is designed to mitigate exploits like this, see my talk from KVM Forum 2018: https://www.youtube.com/watch?v=YAdRf_hwxU8 https://vmsplice.net/~stefan/stefanha-kvm-forum-2018.pdf
Only on platforms that use SELinux though.
QEMU runs on other operating systems like *BSD, macOS, and Windows. It is less mature on those platforms and it's safer to avoid running untrusted VMs on those platforms.
This was in the days when people had dial-in accounts to a shell, and wanted to use Mosaic web browser on their machines.
It creates what looks like a virtual NIC (literally in qemu's case, indirectly by SLIP in the original Slirp), and reassembles the packets it gets coming in from the guest OS or SLIP user. For example, a SYN packet gets turned into a call to connect(), and a data packet gets turned into a write() on the appropriate TCP socket FD. An RST packet gets turned into a call to close(). The reverse happens in the other direction, based on the read() data, fake TCP packets are generated, a closed socket gets mapped to a RST packet, and so forth.
From the outside, it looks like the guest (or SLIP user) is NATed through the host's IP address. But really it is just reassembling the intention of the guest based on the packets it is sending and calling host kernel functions to cause the same effects.
Let's hope that these projects are not affected, too
https://github.com/rootless-containers/slirp4netns/security/...
Also, v0.4.0-beta.2+ can harden its own process by unsharing mount namespace and pivotting_root to an empty dir that only contains /etc and /run with noexec mount option. v0.4.0-beta.4+ additionally supports seccomp filters.
The worst of our offenders was a couple—-a husband and a wife—-who both spent nearly all day on MUDs. One of our employees knew them and told me that they lived in a trailer in squalor. Their kid ended up getting taken away from them by Child Protective Services because of neglect. It was a really bad situation and colored my opinion of hard-core gamers.
> This was especially useful in the 1990s because simple shell accounts were less expensive and/or more widely available than full SLIP/PPP accounts.
While users using Moasic/Netscape on dialup were likely to be pegging their modem link the whole time, and almost all of those bits were bits you needed to buy from a transit provider over an expensive leased line.
Then ISPs didn't like people running SLIRP, so often there were "no SLIRP" rules. :P
1. Exploit a miscalculated pointer to write arbitrary data.
2. Exploit a ASLR infoleak to figure out the target.
3. Use (1) to create a fake timer with a callback to "system()".
Can some forms of Control Flow Integrity mitigate this type of attacks?
W^X is useless in this case, but if Control Flow Integrity offers code pointer target verification, it could have a chance to catch the final bogus callback, am I correct?
Out of interest does Rust prevent this kind of mistake?
Most of what Rust improves on is related to memory ownership and concurrency - we’ve had ways to prevent this class of problem for a long time, and in fact there are even C variants that can do that too.
At a surface level, the bug is in doing raw pointer arithmetic: determining the size of some value by subtracting two pointers from each other, incorrectly assuming that both pointers are within the same object. There's a codepath where one is not, and therefore this size computation is incorrect. Later, that size is added to another pointer, allowing for out-of-bounds access, overwriting other variables.
Rust doesn't let you subtract two pointers from each other. Even unsafe Rust does not; you'd have to cast the pointers to integers, first, because finding the difference between two unrelated pointers is a fundamentally meaningless operation. (Indeed, it's undefined behavior in C, and Rust compiles through LLVM and would inherit the same optimization passes that wish to consider things UB, so it doesn't pass a request through to LLVM that's going to be undefined.) And safe Rust doesn't let you index to an arbitrary spot in an array / buffer without a bounds check, so even if you got a nonsense offset, it would crash instead of overwriting unrelated values.
At a slightly higher level, it seems like the underlying issue here (if I'm reading the article right) is that struct mbuf has two ways of representing the data: the array member m_dat and the pointer m_ext. Which one you're supposed to use is represented by a flag. The code correctly kept track of which one to use in all cases except one. Entirely apart from the memory safety stuff, Rust gives you tagged enums (enums with data, aka "sum types" in functional programming) with the property that you can only access data inside a particular enum variant if the variable you're looking at is actually of that variant. So, for instance, you could have something roughly like:
enum MData {
Internal(buffer: [u8; 32]),
External(ptr: &[u8]),
}
and syntactically there's no way to get a ptr out of an Internal or a buffer out of an External, so you couldn't have the logic confusion that led up to the memory unsafety. Even if you could do raw pointer arithmetic in Rust, you'd still get it right: let delta = match mbuf.data {
Internal(buffer) => q - &buffer,
External(ptr) => q - ptr,
}
so it's impossible to forget to check the flag. (In this case they do check the flag but it sounds like they're not checking the right flag or something? Or the flag is set too early? I don't totally follow the description, but if it's something like that, using a Rust enum would guarantee that the "flag" accurately matches whatever you're looking at.)The memory safety stuff is great, but I really think that having a richer type system like this is more fundamentally what prevents bugs, compared to C where all you have is numbers, pointers, structures, and structures-where-things-overlap. (Another good use of this is nullable pointers that force you to do null checks before dereferencing them, and a little more broadly, this pattern also gives you locked data that forces you to take the lock before dereferencing the data, avoiding issues where you take the wrong lock, which could end up as memory unsafety eventually.)
FWIW there are a few hypervisor projects in the same space as QEMU that are written in Rust: AWS's Firecracker and Chrome's crosvm come to mind.
I know this is only tangentially related to your point, but all the research I have read about points to the opposite conclusion - richer and stricter type systems don't have a proven effect on bugs, whereas memory safety is guaranteed to eliminate whole classes of bugs.
One guess why this would be true, despite your good example of a type system-level fix that would have prevented this bug entirely, is that memory safety is automatic (or at least opt-out), while the richer types are opt-in: nothing in Rust (or Java etc) would have prevented writing the original C, the programmer would have had to think about using the enum or other equivalent in order for the compiler to help. Granted, in this case they would almost certainly have done so, as enums are just so much nicer than flag checking, but for other type-level solutions the same may not happen.
One thing I'm curious about is if there's a slightly different property than "richer type system" that helps. In Rust but not necessarily in other languages with good type systems, you can't leave a struct member uninitialized, so perhaps the easiest way to write this code is actually to use an enum if you know you'll only ever use one or the other. That might be close enough to the "automatic" property?
But Cyclone is a separate language (and so is Rust) because in order for memory or type safety to be useful, you need to carry information that cannot be represented in the C type system. The only information a C pointer carries, besides the pointer address, is the C type. You can't have it carry a lifetime, a bound, or a type that isn't a valid C type. So you lose your safety properties at the boundary between C and Cyclone/Rust. You can't pass a tagged union from Cyclone/Rust into C, because the C compiler isn't going to enforce that C code updates the tag properly - if it did, it would start rejecting valid C code. At that point it's just a matter of style/taste whether you have a language that looks like C but isn't (Cyclone) or looks less like C (Rust).
(Come to think of it, C++ also has the same property of being built from C and being mostly backwards-compatible, but being a separate language.)
Both Rust and Cyclone make it easy to call into C for interoperating with existing code, even though such calls cannot be guaranteed safe at the language level. So it's quite feasible to take a large C project and start converting parts of the code to Rust one at a time, ensuring that you can maintain safety within the parts that you've already converted. (Rust in particular is the unique memory-safe language at the intersection of "actively developed and popular" and "can be used as a drop-in replacement for C and export C-compatible binary interfaces back to any calling code, without overhead" - there are a number of neat languages, not just Cyclone, if you don't have the former requirement, and there are a number of great languages like Go or Java if you don't have the latter requirement.)
If a VM running inside QEMU is able to compromise QEMU's code, it then runs with the privileges of QEMU on the host. The host, the regular OS running QEMU, is colloquially called "the hypervisor." And in most production virtualization setups, there is nothing of interest on the host other than some QEMU processes itself, i.e., anything valuable on the host is perfectly well accessible to QEMU, and therefore to malicious code that has taken over the QEMU process.
The "hardware acceleration" that helps QEMU emulate a computer operates in root VMX and non-root VMX (when actually executing guest instructions) modes. When operating in root VMX mode as a supervisor, it has approximately the highest privileges in the system (ignoring SMM), as it is operating in host ring 0. QEMU runs in host ring 3, making it a typical user mode process. If you manage to compromise QEMU (and only QEMU), you have only escaped into host user mode. In that case you haven't compromised the hypervisor, merely the virtual machine monitor.
If, on the other hand, you manage to compromise KVM (or vhost) you could reasonably claim to have actually compromised the hypervisor.
Type 2 hypervisors make this generally much uglier than type 1, but it's reasonable to say that the hypervisor in a type 2 deployment is limited to the kernel component rather than the user-mode VMM.
and MINIX in case of Intel :-)
In a well configured KVM instance QEMU obeys the principle of least privilege as much as possible, that is it can only access the resources it needs to do its job.
In practical terms, this means QEMU is confined (via SELinux, cgroups, Unix permissions, seccomp) to only access resources for the VM it is running; taking over QEMU does not give you access to "anything valuable on the host".
Of course, having remote code execution in QEMU is awful; it can be a base for exploiting a kernel vulnerability to get root access. Luckily this bug would not be exploitable in a production setting.
And I think that there is no common configuration of OpenStack or libvirt that assigns unique access control labels per VM. Every qemu runs with the same privileges, ergo one exploited qemu can laterally attack another. (But maybe there's something nifty I'm missing?)
(Agree that for the use case of desktop virtualization, a good SELinux config can keep it from accessing your browser cookies.)
There is interest in using it in container runtimes as well, and hopefully this will give slirp more love. It's very old code that most QEMU developers wouldn't have touched with a ten-foot pole...
> Unfortunately, there are no automated tests available.
> You may run QEMU -net user linked with your development version.
Not trying to shame anyone here, but I do think that this would be a good entry point for anyone trying to help this project out. Hell, this might be a fun little target for property testing or other kinds of generative testing