Unikernels: No Longer an Academic Exercise?
250bpm.com
250bpm.com
A few years ago someone trying to spin container technology did a lot of damage to other attempts at unikernels with marketing dogma, not reality. Claims of non-debugability and other FUD. It's long been standard operating procedure to trace and debug systems from outside the system's view. This is how people do bringup of new chips as well as OS ports. On hardware there is usually a dedicated debug and trace facility as part of the CPU or board support package or a firmware monitor. In a virtualized environment like a unikernel this is way easier because you can run code above the guest's ring 0 supervisor privilege against its RAM/pagetable root. Modern systems like POWER even allow debugging from the BMC, with a plan to allow full gdb sessions over that out of band interface https://github.com/open-power/pdbg.
There's nothing implicitly wrong with unikernels, or other systems software ideas like microkernels, just because they are less popular technology at the moment. I'd encourage people to continue exploring this space.
From the perspective of someone who'd debugged and traced unikernels from both inside the runtime (LING) and outside the runtime (xentrace), and who worked in the domain of the very z/TPF system you referenced above, the indictment seemed at-best strangely misguided and at-worst intentionally duplicitous.
As you can imagine the indictment wasn't at all persuasive to me, and thus I keep exploring the space in bits and pieces where applicable.
It's not the corner of the sandbox that I play in any longer.
That said, I never really considered Genode in contrast to unikernels to be honest. Something to chew on certainly.
Which is fine if you're upfront about it. He wasn't.
Now, I'm not sure I agree that I have a horse in the race. I don't necessary believe that there is a race. I've never really been a proponent of the schism between Unikernels and Containers. I struggle to see how Unikernels can offer the same flexibility and ease of deployment as containers. We're likely won't be able to support the vast amounts of runtimes and infrastructure needed to replace something like Docker. Perhaps there could be very specific uses where something like the paper described could be used, but I'm not betting on it.
As a software project IncludeOS has a much narrower target than what people traditionally have thought when thinking of Unikernels. And as a result of of this we're not in the business of replacing neither containers not general purpose operating systems(GPOS). We're aiming to carve out a few niches where we are confident that a GPOS isn't the answer. We're only going to address those needs where we're pretty certain we can actually add some value. Basically we're think we can improve on security in addition to adding real time capability whilst still remaining source-code compatible with Linux (mostly thanks to musl).
My grief is singularly with the myths you helped create that Unikernels are something where you are forces to work with stone age tools and hardly without any tools, except printf, for debugging. We've had to spend a lot of time dispelling these. There are a few other things I believe you where wrong about at the time but I'll spare you the details. Better suited discussions over a beer of coffee.
1st. I don't see that many bugs in the TCP/IP, filesystems, etc. specially those exploitable by black hat hackers, specially remote exploits. We usually see wide-spread vulnerabilities in the middleware like OpenSSH, which is in the userspace, putting everything in the app is not going to solve that, quite the contrary, we need to wait for the library maintainer to fix the bug, for the applications to integrate with the newer version of the library, and then, update all the applications.
2nd. Side-channel attacks. I'm not really confident in letting applications have that low level of hardware access. I mean, it's really easy for one application to steal the resources from another. We need authority if we are sharing peripherals. Well, we have one peripheral already that is shared by processes, that's memory and look at the mess that is a modern MMU and all the attacks on those that we have seen throughout the decades.
3rd. If we are going to implement that as a shared library, the result is exactly the same as today's micro and hybrid kernels, if we are not going to save memory and implement that as a static library, it's exactly the save as virtualization.
Correct me if I'm wrong.
Once developers get into the workflow of writing applixations that can have fuller control of its environment, rarely one would return back to the world of opaque infrastructure and framework.
Here [0] is a good discussion of unikernels (TheRaven knows what he's talking about). ctrl-f for "You absolutely can not do the same thing with a conventional *NIX OS"
I agree that if the advantages boil down to saving a hundred MB of RAM of two, and enabling additional compiler optimisations, it won't be compelling to that many folks.
Taken to the extreme though, we can start spinning up instances 'just in time' to deal with incoming requests. Doing this, you could run (and hopefully pay for) your instances for only a tiny fraction of the time your service is in a working/available state.
This has apparently already been demonstrated [0], but their demo has been broken for as long as I can remember. I presume it worked at one point. I don't know if any commercial Xen-based offerings can spin up an instance as quickly as their configuration did, but the idea is there.
> Explain this to me then: What's the difference, conceptually, between (A) a VM host running a unikernel capable of solving one particular kind of problem, and (B) an OS running on bare metal running a process capable of solving one particular kind of problem?
They are very similar. The main difference is the shape of the OS APIs. In the former scenario, the host provides things that look like CPUs, NICs and things that look like block devices, with no high-level abstractions (e.g. threads, sockets, and filesystems). If these are needed, then the unikernel will link in device libraries that provide them. In contrast, in the later situation the OS will be providing these even when they are not needed and they will add overhead to those code paths.
Additionally, the amount of process state owned by the OS is far greater in the latter case. Things like file descriptors, thread priorities and so on make it far harder to snapshot or migrate an OS process. Hypervisor interfaces tend to be as close to stateless as possible.
To add to this, with the approach we're taking with Solo5 [1], hardware virtualization is only used as one possible sandbox/isolation mechanism. If you look into the code, you'll find that the abstractions presented to the unikernel by the hvt tender are much thinner than those presented to a guest OS by a traditional VMM (e.g. KVM/QEMU).
1. Regarding bugs: There is a lot of code in your monolithic OS kernel. There have been, and will be, a lot of vulnerabilities in that code. Various sandboxing mechanisms notwithstanding, the vast majority of processes running on your system can potentially use all of that code as an attack vector. Unikernels let you switch off access to most of that "host" code, and, through the use of a library OS, contain only the minimal set of libraries needed to run your application in the "guest". Regarding updates: Valid point, but no different from any modern application of substantial complexity, except that now you also have to update (e.g.) the library providing its TCP stack.
2. Unikernels don't make that problem any worse. In fact, if you run on something like Muen (https://muen.sk/), it'll mitigate a bunch of these attacks by giving you a less precise RDTSC at the subject ("VM") level.
3. I don't follow. Implement what?
The paper suggests using system-call filtering via seccomp or similar to block off this attack surface, but if you trust seccomp to have zero bugs, you might as well do traditional process sandboxing a la Google Chrome.
The problem unikernels solve is that traditionally, a vulnerability in C code means "game over" and the box is completely pwned with all memory and storage directly accessible, while with a unikernel, a vulnerability means only that an attacker can access other functionality within the app, and can't rely on eg. a shell or a debugger being available to arbitrarily move bytes around.
Correct. Conceptually the same applies to all Type 2 hypervisors. Type 1 less so, but you could still potentially exploit Xen and you have all of the dom0 to play around in.
> The paper suggests using system-call filtering via seccomp or similar to block off this attack surface, but if you trust seccomp to have zero bugs, you might as well do traditional process sandboxing a la Google Chrome.
If you do that, you will have to do a line-by-line analysis of the code you want to sandbox in order to determine exactly which syscalls it's using. With the approach presented in our paper the developer does not need to care about this. For example, she can develop her MirageOS unikernel as a normal UNIX process and switch to a guaranteed-to-work-minimal-seccomp sandbox with a simple change of target (build-time configuration option).
But yes, you are now trusting seccomp instead of KVM. I believe in giving people the ability to easily make that choice.
> The problem unikernels solve is that traditionally [...] can't rely on eg. a shell or a debugger being available to arbitrarily move bytes around.
That part does not change with the sandboxing mechanism changing. The stuff that's inside is still only your unikernel.
Which do you trust more?
For the 80% case, on x86_64, I consider them more or less equivalent. KVM is used daily in anger to provide isolation (e.g. GCE, and now ChromeOS) and has been around much longer but you need to trust hardware virtualization which is a large attack surface on the CPU itself. Given what we've learned about CPU vulnerabilities over the last year, I wouldn't be surprised to find some lurking in the VT-x/SVM implementations.
Seccomp OTOH is difficult to use correctly for arbitrary/existing applications but exposes less of the kernel (depending on your metric, see our paper) and does not need hardware virtualization.
For the 20% case, where the stakes are higher (e.g. High Assurance), I would use something like Muen or SeL4 and run a disaggregated system on top of that.
I'm aware it's not an entirely accurate metaphor, but it might turn out to be a more accessible one for opening a conversation.
*EDIT: Unikernels are neat as replacements for VMs since the abstraction layer can be much less costly/higher performance. They cannot be used as regular applications running on the OS due to how security is implemented.
This code is usually there for a reason and a unikernel will have to implement large swaths of this code anyway.
Networking isn't trivial and I don't think anyone will be reimplementing the TCP/IP, UDP/IP, ARP, DHCP and more stacks without breaking in a few bugs themselves.
Example, modern TCP congestion control requires a very precise packet pacing and timing.
Another example, TCP retransmission will be required at some point, the code doing this will have to run side-by-side with the app. This will also be some amount of code.
And all this just piles up and up.
The only real advantage I see for unikernels is that because all the hard work is done by the hypervisor they don't have to bother implementing device drivers and task scheduling but end up either replicating or using the hypervisor's own networking stack and a bunch of other subsystems.
The largest part of a monolithic kernel today is devices drivers, by far.
> Networking isn't trivial and I don't think anyone will be reimplementing the TCP/IP, UDP/IP, ARP, DHCP and more stacks without breaking in a few bugs themselves.
Sure. And then more people will use those stacks, and they will get better. The more the merrier, we have too much of a software monoculture anyway.
> The only real advantage I see for unikernels is that because all the hard work is done by the hypervisor they don't have to bother implementing device drivers
This. People continually underestimate the amount of work required to support the hardware ecosystem. This is also why rump kernels (note, not the same term as the unikernel known as Rumprun) are such an achievement, also very much underappreciated.
> The more the merrier, we have too much of a software monoculture anyway.
Linux and some other kernels allow userspace apps to have their own network stack in userspace, latest kernels allow even larger sections of the networking subsystems to be entirely in userspace.
I think this approach should be favored over a unikernel since it uses the natural x86 privilege seperation between userspace and kernel.
There has been exactly 1 remotely exploitable bug and exactly 1 information leak bug in openssh in more than a decade (I believe the exploit is actually more than a decade old, need to check though), it is by far the most secure daemon ever. It is hard to overstate just how much of an accomplishment that is. NOTHING matches it in security. This is even more impressive given how widespread it's usage it.
> We need authority if we are sharing peripherals.
I think you'll find that large websites ... just don't do this. The only real purpose of multiple processes on such webservers is administration.
> f we are not going to save memory and implement that as a static library ...
The reason you want to do this is that the slowest thing in modern processors is context-switching. Going from user-space to kernel space. If you have, say, a networking stack in user space you can avoid such context switching entirely. The speed gains are enormous, especially for webservers.
Ridiculous performance in networking also enables many more applications that just won't work without. What do you think about cluster-wide disks, for instance ? SQLite like databases that run against Petabyte-sized files, concurrently, safely. You're just not going to do that efficiently by having it done in the OS.
2nd. In a pure-unikernel approach, applications can’t steal from each other because there is only one application. Isolation is left to the VM. The goal is to get rid of the redundant isolation of running a multiplexing OS over a multiplexing VM. In the as-processes model, the Linux kernel still provides sandboxing. The goal there is to be a bridge between traditional and unikernel development.
3rd. The goal is to boil down to just virtualization and only the libraries specifically needed by a single specific application.
That said, we're not moving in the direction of unikernels, but it has a clear application in my mind. It's a fair bit of work to migrate, and I've yet to see a practical comparison, so it's unclear what the actual benefits over conceptual benefits would be.
https://github.com/rumpkernel/rumprun-packages/tree/master/e...
http://www.erlang-factory.com/static/upload/media/1474729921...
I'm not sure if there's any news on it since then.
1) (Putting on my black hat) Attackers don't care about bugs or exploits. They care about running their code on your system. Whether that is as simple as mysqldump or wget'ng a monero cryptominer to run on there it all is based on the premise that the monolithic operating system (whose design is ~50 years old and linux is > 27 years old) is explicitly designed to run multiple programs by multiple users. Keep in mind this design pre-dates commercialized virtualization (eg: vmware) and pre-date "cloud" (eg: aws). If we assume that you are already utilizing VMs (and you are if you are on any public cloud and you are in most private on-prem deployments) the VM then becomes your isolation model. Can you still attack the underlying infrastructure? Sure - but if you can root GCE or AWS I'd say we all have some serious thinking to do on the current state of cloud infrastructure. Contrast and compare that to all the ridiculous headlines you see every single day and the fact that every single RCE that is worth doing entails forking/execve a new process. It's one thing to have the instruction pointer - it's quite another to launch a shell that doesn't exist, link your program to libraries that don't exist, as a user that doesn't exist, download new code when you can't.... etc.
2) Not to belabor this point but side-channel attacks affect everyone and Intel has been taking the hard (in terms of market) approach of simply disabling hyper-threading on some of their hardware.
Security is the number one selling point for unikernels imo.
See, for example: https://www.kernel.org/doc/html/v4.11/gpu/i915.html#batchbuf...
See r600_packet0_check, r600_packet3_check: https://github.com/torvalds/linux/blob/master/drivers/gpu/dr...
NICs are an unfair comparison because the commands for NICs are all generated by the kernel, which is trusted. For GPUs, commands are generated by an untrusted user space driver that translates OpenGL or Vulkan commands into command packets for the GPU. Those command packets contain memory addresses, so in order to achieve isolation between different user space processes in the face of a potentially malicious user space driver, the kernel used to have to validate those addresses in the absence of GPU page tables.
Anyway, this kind of validation really hasn't been necessary for a long time now, despite pcwalton's outdated information, because GPU page tables were phased in around 2010. (The first AMD chip to have them was Cayman in 2010. Since we're citing kernel drivers here, take a look at https://github.com/torvalds/linux/blob/44786880df196a4200c17... and note how the function skips any parsing when virtual memory is used.)
Re: 2 - this is managed by the hypervisor. SEL4 can be run as a hypervisor and has been formally verified to be correct [1]. Trusting a proven system like SEL4 is a far-cry from trusting Linux's isolation primitives because we can make hard-guarantees about the behavior.
Unikernels have the advantage that they are basically backwards-compatible (you can run any VM on the infrastructure you develop, even if that infrastructure is tuned for unikernel-native applications). With Unikernels, you can achieve VMs that are lighter than containers [2] thereby increasing your customer-per-server ratio.
[1] https://sel4.systems/Info/FAQ/proof.pml [2] http://cnp.neclab.eu/projects/lightvm/lightvm.pdf
> Correct me if I'm wrong.
OpenSSH is probably the single most secure piece of C network software. It isn't exactly a good example for "wide-spread vulnerabilities in middleware".
Also - there are production deployments out there.
https://arstechnica.com/information-technology/2016/12/how-a...
It is also how they kind of implement secure kernel in the recent versions of Windows 10.
So in the server less future, Linux and FreeBSD will no longer be at the heart of these "Functions"? But instead they all migrate to Unikernel? At the scale they are operating I am pretty sure it make sense, if it brings 5 - 10% performance improvement along with other benefits.
Not sure if I like the way things are moving in that direction though.
Then there are those few use cases where one actually needs to dive into OS specific syscalls or Assembly.
So for a large class of applications it doesn't matter if the application is running bare metal, on a VM, container or plain old OS process.
It was a more of a workshop in relation to ICS2018.
https://arstechnica.com/information-technology/2016/12/how-a...
I'm guessing this is mostly interesting for FaaS platforms or similar. You get isolation similar to hardware-assisted vms but with a lot less overhead and with phenomenal boot-times.
That is not to say the approach is not interesting, but it'll probably be for in some kind virtualization, where you abstract away the OS entirely, and run applications directly on the hypervisor.
Default settings always not working for all apps.
If you need to tune network, if that's a tuning for all apps, development can be done inside net lib, all apps rebuilds and release, that's not more expensive than tuning a kernel config.
If you need to tune on per app basis, then unikernel is wildly safer because of the isolation.
Just use Linux or something similar and use the lower APIs.
E.g. don't mount a filesystem, use the block device directly in your app.
There are others as well, but AFAIK these are the ones with active development happening. Please correct me if there are active ones I've forgotten.
All three support compiling your application as a Linux binary, meaning you can do most of the development and debugging using the tools you're used to, IDE with visual debuggers and the whole shebang.
Once you want to run in a separate VM you can still debug, but it becomes a tad bit harder. I know how to debug IncludeOS application when they are running under Qemu or KVM. Qemu can act as a gdb remote, you just need to get it working, which is a bit of pain.
And now you have the option of running it the way it is describes in the paper. I've never touched this so I don't have a feeling for how hard it is getting it to run.
https://www.joyent.com/blog/unikernels-are-unfit-for-product...
Bryan points out many weaknesses of unikernels, but assumes everybody needs multi-tenancy and the ability to spawn processes. And debugability is weak, but that probably depends on the application / environment you run. A lot of people want to run unikernels in a VM environment, but for me, it seems like the right application of unikernels is where you have one application that you want to expand to fill a single machine -- bare metal, boot to the application, save all the layers; there's no multi-tenancy, but if I'm running on hundreds/thousands of machines, I don't need multi-tenancy.
Debugability is important, but lots of people run without a kernel debugger, so it might not be that important to everyone. If you want it, you'll have to build it, but DTrace and friends had to be built too -- and it's easier to build it a second time, since you know it's possible and what it should look like.
There's a lot more out there - unikernels in particular are similar in nature to RTOS (like the system that is on the Mars Rover) and they definitely descend from the microkernel branch of the tree.
Having said that many microkernels tend to be multi-process. There are now over 10 different unikernel implementations out there currently. They all have different aims and goals but I'd say one defining characteristic amongst all of them is that they are single process by design. There's many other considerations but that's the most important one that comes to my mind.
Microkernels are about splitting a single OS into what are sometimes called 'servers', whereas unikernels are individual systems running on a hypervisor; ultimately, the biggest advantage is that, in a microkernel, if one 'server' goes down (the disk server, for example) the whole system is down, whereas a single unikernel can fail without impacting any other unikernel.
https://jdebp.eu/FGA/microkernel-conceptual-problems.html
https://utcc.utoronto.ca/~cks/space/blog/tech/HypervisorVsMi...
Rather, if the disk server goes down, the whole system might go down. It doesn't necessarily will go down. You could, for example, simply restart it. If your Graphics server goes down, you might still be able to SSH to the machine, and in the example of yours which I quote you could run a remote emergency SSH server in full RAM.
Unikernel and microkernel are good examples of (fine-grained) principle of least privilege. Doing that right is difficult. Just ask the NSA or Intel.
1. The disk hardware is bad. By all means, take everything down until you replace the disk so you don't get garbage in important files!
2. The disk driver software is bad. By all means, take everything down until you replace the software so you don't get garbage in important files!
In neither case is simply bouncing the disk server an acceptable answer. Debugging and fixing the underlying problem is essential, and that requires taking the whole system down.
Probably #1 (or a nefarious person being root), but you do have backups and redundant servers, right? I'd read SMART logs first. Back when my Death Star died some 20 years ago I remember it getting more difficult by the day to spin it up. But once it worked, it'd work fine.. until the machine tried to read from a bad sector (a good OS wouldn't have frozen on that but this was Windows 9x).
Here are some alternative possibilities:
3. Physical compromise; e.g. a cable got yanked.
4. Disable write caching and continue despite the driver or hardware being bad.
5. There is no quick way to get physical access.
6. A disk server isn't necessary to keep the system up. Goes well with #4.
If you look at how it all went, the massive redundancy of cheap commodity hardware won. Containers add further on that, and no doubt more bloat can be removed via a unikernel or microkernel (leading to smaller attack service but also more complexity). But that doesn't mean every aspect of such a machine needs to keep running, and it does imply a FOSS solution.
I've had times with bad sectors (Deathstar at the very least) where reads would result in complete OS lock up or very laggy situation. That can be solved by only trying so many times to read (though that was on Windows 9x with PATA).