After Docker: Unikernels and Immutable Infrastructure
medium.com
medium.com
Will be giving a talk on this at http://operatingsystems.io/ in London on November 25th.
You're right that the documentation for the rumprun-* stacks (running unmodified applications on rump kernels + Xen/bare metal) is lacking, however, it's early days for this work. New, extra fresh, hot off the press! We're working on it :-)
Note that rump kernels as such (sans the Xen and baremetal stacks) are well-proven and tested code, perfectly usable in production today.
If that is lacking, please let us know!
IIRC Linode (or some similar company) used to use User Mode Linux, but switched to Xen for performance reasons.
see: http://wiki.rumpkernel.org/Info%3A-Comparison-of-rump-kernel...
Also rump is BSD licensed, Grapheme is GPL, which is useful as you have to link it to your application.
The ZRT[ZeroVM RunTime] also replaces C date and time functions such as time to give program a fixed and deterministic environment. With fixed inputs, every execution is guaranteed to give the same result. Even non-functional programs become deterministic in this restricted environment. This makes programs easier to debug since their behavior is fixed.
I've had a play with it - there's a version of python that runs on it, and it's surprisingly usable.
At work, our tech team has found an interesting way around this for our Python app. We build out the virtualenv in the docker container, and then run our ansible-based deployments inside the same container. With that, our virtual environments are rsync'd to the app servers so we can avoid installing developer tools.
In the context of Docker it's inconvenient, but entirely straightforward to create images that don't include those elements.
It's a matter of creating an unwieldy chain of build steps to avoid committing intermediate containers.
If/when something like this [1] gets merged things will be greatly simplified.
You can make that cleaner by making the build steps into one Docker image, the final app into a second, and have them share a base image that contains all the basic dependencies.
For Ruby at least the intermediate build step would typically only need to be re-run whenever your Gemfile/Gemfile.lock changes.
For VMs/containers that already run a single application, except for some weird edge cases, there's really no point in having a virtual environment in a virtual environment.
I have initial success with a few simpler projects, now looking into transitioning more complex ones. Not sure whenever it'll go without any hassle, but seems worth trying. At worst, I'd just waste my time and return to virtualenvs.
https://github.com/spotify/dh-virtualenv
One reason to keep virtualenvs is that the system Python (VM or container) includes extra Python packages that your app may or may not need. If you use a virtualenv, you exclude these system-installed packages and guarantee a clean starting point.
I'm running a script right now that generates an ISO that turns a brand new machine into a server running our app with a template DB in completely unattended fashion.
Seems to work pretty well, although I haven't tried it for anything production related.
The idea of the unikernel and the libOS in general where applications can be linked with their bare minimum OS runtime and packaged is certainly nifty, but it's kind of funny that people are being so hyped over what sounds like a more advanced form of what was regularly done in mainframes 60 years ago.
If the programming language has a rich ecosystem with a runtime that is already taking care of hardware abstractions and scheduling, why replicate it a few times in lower layers?
How many schedulers or device drivers are needed to serve network requests?
I'm talking more about the mainframe level stuff, like z/TPF on the software side and adding address space tags to the channels on the hardware side. Basically that last one is a better implementation of an IO/MMU where a device knows that it's probably running under a hypervisor, can get requests directly from multiple VMs without the hypervisor's intervention, and the VMs requests are implicitly tagged with the address space of the VM so it's still memory safe. ie. a VM can't request DMA from a device that would be outside of it's allocated space, but it can still directly ask the device for DMA without involving the hypervisor directly, and the device can service multiple VMs (that last bit is what isn't really present with current IO/MMUs).
I was quite surprised to discover that Java, .NET, Android, Windows Phone concepts were already successfully implemented in the marked in such systems, so many years before. And that the majority of developers out there are unaware of them.
Thanks for the z/TPF overview.
As for lack of VM resizing, that is a hypervisor/Unikernel implementation detail.
If you want to see what happens when you aim for a single-purpose isolated unikernel design, but then bake it fully for operational requirements, look at an (embedded release package of an) Erlang application. There's still a lot of "stuff" there—a lot of attack surface that has nothing to do with achieving the purpose of your app per se—but it's all necessary to keeping your app healthy and stable in the greater ecosystem of services it interacts with.
(This is presuming that you can't just shunt off these responsibilities to the hypervisor. If logging means "your unikernel writes to the console and Xen pipes it to rsyslog" then a lot of problems do go away.)
As a proof of concept, several months ago I built a few tiny Docker images using musl libc and no package manager. But I had do deviate from the normal image build process to do so.
I personally played with aboriginal Linux, but I believe it's the same idea :)
Could you clarify what you mean here? Is this specifically a Unikernel problem or an ecosystem problem (in terms of actually trying to deploy Unikernels in the wild)? If so, those seem like different issues and should be discussed separately.
On a technical level, VM memory hotplug is probably necessarily slower and flakier (ACPI anyone?) than changing one setting in a cgroup.
Sure you could optimize the heck out of the hypervisor, but now you've created a kernel. And your applications run on that kernel.
With containers, you have one kernel that won't have to instantiate 20 drivers for the disk subsystem. It can be smarter because it knows more about the loads. It's what kernels have been built to do since day 0.
My main concern with unikernels is that eventually the hypervisor will need to be a kernel to be any more optimized. I just worry it will be come something of a self-defeating concept.
hypervisor -> monolithic kernel -> containers -> application
Unikernels collapse it to: hypervisor -> unikernel/application
It's certainly more elegant, although I'm skeptical of the purported performance gains as well, simply because so many optimizations have been thrown into traditional kernels.https://www.usenix.org/conference/osdi14/technical-sessions/...
We describe the hardware and software changes needed to take advantage of this new abstraction, and we illustrate its power by showing improvements of 2-5 in latency and 9 in throughput for a popular persistent NoSQL store relative to a well-tuned Linux implementation.
That said, a simple application like memcached might be currently latency-bound by the kernel's network stack, but a more complex application that reads from disk (even SSD) won't be.
https://github.com/siemens/jailhouse
Since the guest unikernel isn't a full kernel, the hypervisor interface is much more minimal, and the few host features it needs can be delegated to the CPU via VT-X (e.g. page table mapping).
At least, that's the dream. (I've never actually used Jailhouse or tried any of the research projects attempting this.)
I'm not disagreeing on the principle of immutable servers, but that's a pretty bold claim.
I don't see getting rid of the "accumulated cruft" as being a particularly interesting reason for exploring the unikernel or immutable server concept. The benefit is in building for scale and redundancy. The lighter your image, the easier it becomes to replicate it and maintain it, generally speaking.
Further, there is an argument to be made that building your own cruft into your system is counter-productive compared to letting the operating system cruft handle it. The Linux developers are probably better at it than you or me; unless we understand our usage patterns dramatically better and the options for optimizing them, it may be best to trust the OS "cruft" to do the right thing.
In short, I'm not really taking a side on this one. I believe there is interesting research to be done, and probably useful outcomes to be found, in this direction. But, why dismiss 40 years of operating system refinement by some of the brightest minds in the world as "accumulated cruft...of bad ideas"?
I think it's a mistake to conflate path dependence and correctness.
Well I learned something...
It's also largely irrelevant, because CoreOS should in practice be read-only when you boot it, and you're not extremely concerned with the details of how it's put together (which its usage of Portage is).
I also don't find the "problems" with Docker overly problematic.
* The use of many images is probably(?) not an issue? Do people just use "any old base image" without further thought? * An image of a few hundred megabytes isn't small, but it's not terribly large either.
Lastly, I see people's confusing over what CoreOS is besides the point. What it is becomes pretty apparent after taking a look at coreos.com.
Overall I really like the idea of an immutable server though!
So instead of updating packages and what not, you rely on the developer to update the libraries and reship.
Sure its not far from today's model if the dev has to ship the whole container, but it also makes it even harder. How do you know if you have lib x or z when it's sometimes just dropped among a bunch of files? I think it's much worse. it hides the problem and makes it difficult to detect.
I'm suspecting kernels will slowly converge toward plan9-like functionality instead. It makes more sense. It's faster, more efficient, simpler.
The main barrier so far has been portability - but with more and more apps being written on very portable languages (python, Go, C#, ...) its becoming easier.
Um, not it's not.
Heroku's buildpacks code caches a hell of a lot of stuff on each execution agent. Still more code has to recognise and try to repair various broken states. It's mutability, through and through.