IncludeOS: Run your C++ code directly on virtual hardware
includeos.org
includeos.org
The idea behind unikernels is that now that we have pervasive virtualization, most virtual machines are running only a single application anyway. So you have an instance type for memcached, one for postgres, one for Node, one for nginx, one for Redis, one for your JVM, etc. In many cases, these apps already include user-level scheduling and resource management, which often works at cross purposes with the OS. In other cases, they could run much faster if they implemented these specialized to their particular workload, and didn't need to run through an OS scheduler that also needs to support responsive desktop apps.
Though I suppose the solution in that scenario is "Take it out of the unikernel, then."
Still, if you rip out the fundamental abstractions of an OS, how would you run an application that relies on e.g. fork? Or any application that runs a shell command in any way?
My understanding is probably falling short. I should read up more on unikernels.
You probably don't (at least not without removing/changing that part of the application). Unikernels explicitly do not aim at allowing to run a random preexisting application and trade that for the freedom from existing APIs and conventions.
It's important to understand the context behind the current unikernel hype: it assumes that you've embraced cloud providers (or at least run your own cloud on Xen or kvm), it assumes that you are building a distributed system, and it assumes your distributed system is built using standardized components. If 99% of software is various combinations of nginx, Node, Python, Ruby, the JVM, memcached, redis, PostGres, MySQL, and MongoDB, it can make sense to modify just those 10 applications to run on bare metal, cut Linux out of it entirely, realize gains of up to 10x in performance (based on the original MIT exokernel papers), and still present the same programming interface to application-level programmers.
Unikernels are not useful to people who just want to hack C on a single box, nor are they useful to giant companies like Google where all of their software is written in C++ with POSIX APIs. But they could save a lot of money for many mid-range businesses who currently host on AWS or GCE and write largely in high-level languages.
realize gains of up to 10x in performance (based on the original MIT exokernel papers)
Is that true? Do the observed improvements in the field get anywhere close to that theoretical boost?
If that's true, then that's incredible. Which layer of the OS is responsible for an Nx slowdown? Where N is whatever the real multiplier actually is most of the time.
(Aside: it's probably not a good idea to imply "high level language == doesn't use fork." Fork is a fundamental primitive. There are things you can do by forking which you can't do by other means. I mean "can't" in the same say as "yes, every language is turing complete and so therefore can simulate any other language, but you can't write recursion in BASIC, because you wouldn't want to try.")
It just so happens that it feels that way because it's so universal in multi-process systems. The reality is that there are many uni-process systems, like embedded microcontrollers etc.
Unikernels actually have a lot in common with these sorts of environments. For instance if you are running a unikernel on PV Xen or KVM you have direct access to the paravirtualized network and block device buffers, avoiding additional context switches when doing IO (this is where those 10x speedups from from).
Thing is though you can achieve the same performance on more traditional OS + multi-process setups too but it does require some level of device level virtualisation + isolation and kernel bypass. For instance combining PCI SRIOV with say PF_RING would give you effectively the same performance profile without sacrificing OS level features.
https://pdos.csail.mit.edu/exo.html
The speedup they observed was because they could bypass all the layers of the OS and implement abstractions specialized to their particular use-case, making use of information that is available to the application but not to the kernel when deciding how to allocate resources. For example, Cheetah stored preformatted IP packets on disk, which would be sent straight out to the NIC along with the file contents from the filesystem cache. There was no need for the overhead of a TCP/IP stack, no need for buffer copies between kernel and user mode, no need for scheduler overhead as the process is put on the wait queue while waiting for the filesystem, no need for kernel-mode context switches.
I could imagine several similar cases with modern apps, particularly as the NUMA penalty has grown and we've gotten hardware technologies like RDMA. Imagine a Node.js implementation with a locality-aware scheduler, for example: instead of running whichever closure is attached to the file descriptor that epoll happens to return, it preferentially executes the closure that was most recently enqueued, on the theory that all of its context is likely to be hot in cache. Or imagine memcached with full control over the TLB, so that you could lock certain entries (eg. active session objects) into the TLB and never page fault on them.
Obviously this is not suitable for every programmer (e.g. if you have difficulty designing software that does not leak memory, this is not for you) but there is a subset that can design correct schedulers and robust resource managers in their sleep. The integer factor improvements in system throughput make it a worthwhile optimization if you know how to do it.
I think the operating system for cloud services is need to be able to run heterogeneous system. Which Mesos & CoreOS are headed towards this idea via containerization.
Why build an OS that only run a C++ for cloud services? Is there any use case/problems that the author trying to solve?
Edit: looks like there's a direct alternative called Rumprun: https://github.com/rumpkernel/rumprun. I hadn't heard of it.
Anyway, right now this whole unikernel thing is still pretty new so it's good to have multiple projects.
While this may be the most typical use for OSv, my understanding from their wiki[0] is that OSv can handle Ruby or Node, or Linux apps (most of the Linux ABI is supported).
smartos zones also has a linux compatibilty layer ( lx branded zones ).
It's a 1950s idea updated to 21st century standards using 1970s dynamic linking, sometimes but not necessarily meant to be run under a hypervisor.
(And see the demo here: http://zerg.erlangonxen.org )