Unikernels: The Next Stage of Linux's Dominance [pdf]
cs.bu.edu
cs.bu.edu
Edit: In case you can't read the paper there's a copy here: https://www.cs.bu.edu/~jappavoo/Resources/Papers/unikernel-h... (Thanks anonymousDan in the comments below for linking to it)
Since last HN post about GNU Hurd, I'm wondering why not try to join this Free microkernel project instead of looking in another kernel's direction ?
Unikernels strive to minimize communication complexity (and size and footprint but that is not relevant to this discussion) by putting everything in the same address space. This gives them many advantages among which performance is often mentioned but ease of development is IMHO equally important.
However, the two are not mutually exclusive. Unikernels often run on top of microkernels or hypervisors.
While a unikernel might be structured to some extent internally into modules, it's basically all linked into a single blob and running in a single address space. Some programming languages may support some kind of PL-level separation, but there is no hardware enforcement. In the case of what Ali is doing because all the code is written in C (it's all Linux, glibc and memcached) there is neither software nor hardware separation internally.
Spin OS (Modula-3), Singularity (Sing#), Midori (M#), TockOS (safe rust).
As far as I know TockOS is the only one I know of which has some form of both PL-level and hardware enforcement of separation, albeit on an MPU rather than a full MMU, PL-level for kernel modules, and an MPU protected userspace.
I at least think it is worth addressing that none of these separation mechanisms are actually mutually exclusive.
Strip, as a UNIX tool removes the textual crap in a debuggable binary, to leave behind only the bits you needed.
It feels to me like if you can do runtime code call checks, and confirm which actual calls you make, then stripping the bits of libc and associated libraries out, being left with only the strictly required calls, and then by extension the syscalls, and then by extension the kernel elements, is actually possible a lot of the time.
So, library -> reduced library -> reduced calls -> reduced syscalls -> reduced kernel state is a sequence or set or something, of applied minimisations which can be done, if you can predict all the call paths in your code and their dependencies.
But then its not a generic OS any more: its an application specific binding to a general purpose CPU.
Why not go all the way, and work into the ALU and remove the bits you don't need? And then go into the micro code, and the associated FPGA, and everything else..
Edit: more informative link https://fgiesen.wordpress.com/2012/04/08/metaprogramming-for...
So yes, I think you're right: great dream, hugely hard to do, if not actually impossible in most cases near as damnit, but there are places where program proofs take you to this, Military/Space/Medical needs to know the code calls don't have unexpected outcomes. And language choices which express as "how" you express code, can also help. Probably not enough.
[0]Type-Driven Development with Idris
On the other hand there is tremendous value in using some well-defined standards along the way so you can benefit from all the debugging, analysis and verification tools that already exist, so such hardware/software co-design is always a balance.
Because there are diminishing returns...
get a runtime profile, then optimize the code based on how it will run.
linktime optimization has the compiler output bytecode files, then the optimization happens at the link step where it happens globally across all objects.
If yes, does anyone know if there is/will be any automated way say to analyze a software, example: a .deb package containing an user software, then build a list of dependencies and create a bootable image for the requested architecture which will run just that software?
In the past I've brought up Linux on embedded boards with 4MB RAM and 4MB flash for both the kernel and the application, and we gained most by linking statically to a smaller libc on the user side and stripping out drivers on the kernel side, ditching bash for ash (currently dash) and using a stripped down barebones init. We could have ditches ash and init as well and booted straight into the app, but ash made debugging and testing slightly easier and we could fit it in. 4MB at the time was not even a challenge.
I think the gap from something that you can fit a Linux kernel and an app mashed together into a unikernel up to being able to run them separately will be very small, so in most instances if you can make it work as a unikernel odds are you can produce a nearly a small solution without it.
Going full on unikernel potentially costs you a lot of resilience against failure, it's not just about security. The main reason for wanting to do it is to reduce the amount of context switches and data copying to reduce latency.
So does the normal kernel build process once you've specified which drivers you need compiled in. That's in no way unique to a unikernel approach. There'll be extremely little code that can be stripped from a unikernel that can not be stripped from a kernel built for the purpose of running a specific application.
It's just not where the benefits of a unikernel are. As I said, and you reiterated, it's really cutting latency from syscalls that is the big benefit of a unikernel.
If a few MB of space matters, then you'll want a suitably configured kernel that disables a ton of stuff that is totally irrelevant for such a system, and you'll need to take that configuration into account whether or not you're building a normal kernel or a unikernel, as it involves not just excluding drivers and the like, but disabling functionality that if left enabled will not get removed by the linker because there will be calls into it from code paths that are reachable.
Do that, and you'll find it is pretty straightforward to get a Linux kernel down under 1MB.
Now, it's possible your tooling makes that easier, but my point is that it is equally possible to disable those parts of the kernel for a non-unikernel system as well, and you're not going to make significant additional savings. And if you want to, then ditching glibc for one of the smaller libcs would be a better starting point.
This is not a criticism of this project - size is simply not what most people look to unikernels to address, especially not one built on a general purpose kernel.
Building minimal bootable to be very small is possible, but is difficult in comparison to an apt install - I imagine for the most part porting docker container configs to a bootable OS might be the best approach for small-medium projects.
The problem with this approach is that you can make it as small as you want but it's still Linux. At a certain point are you going to start patching things out like support for users? Support for management of multiple processes? There's a non-trivial set of syscalls and data structures designed solely for these constructs. You can't just seccomp it and call it a day.
For us it's not about the size (that's nice of course) but it's more about the performance and security.
I'm not suggesting it's a good idea, but it's there. I'm sure there's more minimal, and less minimal options available.
I don't think there's any security impacts with using alpine Linux specifically, aside from default credentials in a bunch of containers a few months back.
The project you want is "Yocoto", a complete toolchain to build customized embedded OS images.
You would have to build your own analyzer. Grepping the dependencies from the makefile/build data or just parsing the output of dpkg and translating that into yocto build specs is not unreasonable.
Where are unikernels widely used?
To me, unikernels feel quite a bit like statically linked server binaries running under an unprivileged UID - but you're choosing not to trust Linux' (or any other kernel's) user separation facilities, but your hypervisor's domU separation facilities instead. In exchange, you lose virtually all of your existing OS's amazing debugging and performance analysis/tuning tools. It's not a tradeoff I'd readily consider.
That is a simplified view, and there are other advantages as well: you probably don't need to exit to your hypervisor quite as often as you do with syscalls, so you benefit from reduced overhead. One disadvantage is that memory allocation is almost static (unless you do ballooning which is tricky to get right), so you "waste" some memory compared to a bare-metal deployment where unused memory can be used for the page cache instead.
Tooling/Perf/etc is not needed when running (do you really want to debug in production) but tooling can be used in the process of development.
Why consider unikernels? 1) they’re stupid fast. You have only the things you need (booting in nanoseconds? yepp) 2) small attack surface (because you only have your app that’s the only attack surface. you don’t have cruft that build up in you os/kernel over years) 3) light resource usage (you could run thousands of these on the same physical machine) 4) true isolation via the hypervisor
definitely worth keeping an eye on the developments in this space
Whether or not you want to debug in production, reality often means that you will see things in a live environment that you will not see in other environments.
Unikernels are very interesting and have a number of compelling attributes, but let's not pretend that the current state of available tooling for troubleshooting, instrumentation, and general debugging isn't a challenge.
so I actually believe there is an opportunity here to focus on the important pieces (network messages, control flow tracing, memory footprints, etc) after ejecting a huge amount of irrelevant stuff
Definitely excited to see how the technology evolves over the next few years. It hasn't moved as fast as I'd have expected over the last 2-3 years but I'd love to see that accelerate.
Do I want to debug in production? No.
Do I have to do it anyways? All the fucking time. I'm not perfect, I sometimes ship bugs, and when they show up in production, I need to diagnose, determine if rolling back will solve them, or whether I need to fix forward and how. Not being able to debug in production is simply unacceptable.
I want to ship perfect code. But I don't. So, instead, I debug. If the issue shows up in production, I debug in production.
If I can't attach a profiler to find, for example, the code that contains a regex that's doing too much backtracking with production data, then I'm dead in the water. If I can't get a sample out of a system that's received a query of death, then I'm dead in the water. If I can't attach dtrace (or the equivalent) and get stats on how much I/O a system is doing in response to various events, I'm dead in the water.
Being able to dive into the depths of a system is the #1 criteria for confidently being able to put it into production.
I don’t follow. Are you saying that development should build normal user space binaries, and you should only build the unikernel for production?
You will inevitably run into a behavior difference between the unikernel and user space outputs. Even if you’re not debugging in production (shudder) you need to be able to debug the unikernel.
The point (for me at least) is discarding features that you don't need in return for performance and simplicity gains. Those gains are not insubstantial when you're scaling your system and you don't want to spend your money on admin staff.
If you want a shell, that is what stuff like JMX, or server side REPLs are for.
"Lisp at NASA: using a REPL to debug Deep Space 1 while it was 100 million miles away"
The original unikernel paper used a streaming media device as an example and the numbers they provided were really good.
Or the Docker TCP/IP stack on macOS, using parts of MirageOS.
Basically, you've taken the whole /sbin/init link out of the OS init chain, and replaced it with... A single binary you may remember.
To do this took a little extra work. Your glibc and any other dependency ended up in kernel space.
Single-use virtual machines obviously have the biggest boost, but other uses exist. Your kerberos-daemon isolated into its own vm on the network. Someone's gotta serve network filesystems. Apache servers for various websites you may host - the possibilities are endless.
The processor itself has a Massive timesink for isolating kernel-space from user-space (even root is just another user at this level, I'm afraid). Even Server-side admins have simply gotten used to it - every application, server, program Ever runs on that user-space.
So you have a bunch of single use VMs... and that much "context-switching" between virtual-drivers and memory allocation on one side, and the actual server the vm is for on the other. Now these VMs are running on an OS's hypervisor with the same context-switching performance problem!
So, we package each vm as a unikernel - shove each virtualized server into the kernel. We avoid that context switching.
Build that kernel lean enough, and sometimes the unikernel VM can perform almost at baremetal speeds.
What if we did the same for the Host OS - built a hypervisor Unikernel to delegate hardware and nothing else? zomg folks are squeeing with excitement about a new way to frame building a lean system...
But it gets us compiling in-house again instead of using COTS (commercial, off-the-shelf software). Got an old Gentoo-user hanging around that can debug kernel code, harden the kernel against attack, and compile it all lean?
That guy is who we're gonna need to figure out what goes wrong with any of this. Because Murphy's law is unreliably unreliable, yet ultimately absolute.
In the general case, you can almost always download the article from scihub by searching for the DOI of the publication.
OT: This should honestly be a criminal offense, at least if done before a year passes. The publisher does valuable work in vetting contributions and checking the paper before publishing. They shouldn't be cheated out of their fair share of money that comes with due process just because authors go rogue.
But determining whether or not authors have "gone rogue" depends on their agreement with the publisher. You do not have a basis for claiming they have done so unless you know whether or not they have obtained permission or whether or not they never handed over rights in the first place.
Unikernel plus: A lot of code, including code with vulnerabilities, can disappear. So some vulnerabilities disappear.
Unikernel minus: The unikernel and application don't have any security isolation, since the kernel is essentially linked into the appl. As a result, the application has all the privileges of its virtual machine (not the subset that would normally be imposed by the kernel), which would normally be more privileges than it really needs. So any remaining vulnerabilities can have a more devastating effect.
That trade-off in terms of security is really hard to evaluate.
But the kinds of performance improvement shown here, with relatively modest changes, is a really really big deal. So I'd expect a lot of people to investigate this further; it certainly seems promising.
There is a downside, if there is a vulnerability the exploit can probably make hypercalls straight to the hypervisor, but a hypervisor can have less of an attack surface than a full OS.
NCC group did a good article on unikernels and how crap they are if you want to know.
A better idea, like already suggested would be to create an OS which builds itself according to a target application and system, to host that application on said system specifically. that would reduce about as much code, but keep potential security mechanisms like aslr, stack protection , user / kernel separation etc. in tact. (now kernels / oses build to target system ,but not application! -> application would only use subset of kernel, and thus kernel can be built to target application, reducing kernel to whats needed).
don't try to be cheap for performance and skip security, we're not in the damned 80s anymore.
/endrant
>[..] A better would be to host that application and reduce the kernel to whats needed
This is a unikernel.
[0] (large PDF) https://www.nccgroup.trust/globalassets/our-research/us/whit...
https://nanovms.com/dev/tutorials/assessing-unikernel-securi...
I hope people will only use unikernels with memory-safe languages or otherwise well sandboxed runtimes.
The best part of this joke was that I can't tell if it's shade or earnest appreciation.
The link the PDF contains IDs and stuff which differ each time I visit, which seems unnecessary.
Maybe get it from here: https://www.reddit.com/r/programming/comments/cb17mn/read_a_...
which links to: https://www.cs.bu.edu/~jappavoo/Resources/Papers/unikernel-h...
70%-90% of bugs seem to be memory-related. Let's end that with a little bit of effort upfront for much fewer headaches in the long-term.
https://twitter.com/LazyFishBarrel/status/115341007092920320...
https://www.zdnet.com/article/microsoft-70-percent-of-all-se...
https://security.googleblog.com/2019/05/queue-hardening-enha...
Edit: I don't mean to say you have to use C for the "userspace" part of this. It should be possible -- in future -- to use any language for that. However at the end of the day you'll still be linking that with Linux (written in C) and glibc into a single binary that runs in one address space.
Everything looks great on surface but as soon as you start doing low level coding a lot of issues pop up and you need the nightly compiler and xargo and some undocumented library for the boot sequence and whatnot.
TL;DR: great idea. Let's just wait 2 years
Maybe "work in progress" better describes the situation. We need to wait for the language to stabilize, then the tooling and finally the libraries before starting a huge project like writing an OS.