ZeroVM: Smaller, Lighter, Faster
rackspace.com
rackspace.com
The original headline is preserved, and clarified by the editorial clarification in square brackets. It would be great for HN to adopt this as a solution to the modified headline problem, with the provisio that editorial comment must only be used for the purposes of clarification.
ZeroVM: Smaller, Lighter, Faster [RackSpace acquires LiteStack]
How much did you guys pay for them? :p
* Is ZeroVM capable of running unmodified Linux binaries? If not, what compiler toolchain is required to get it working? The main advantage of other lightweight virtualization solutions (OpenVZ, LXC) is that it's very easy to take regular binaries (e.g. postgresql) and drop them in a sandbox with minimal fuss.
- Binaries need to be recompiled. There are two toolchains, a GCC-based one and an LLVM-based one. We can also compile within the ZeroVM container itself.
We expect that a lot of people will use existing language runtimes (Python, Lua, JS) To avoid compilation.
Over the long term, though, a lot of the power comes from composability. Think Unix pipes, in parallel, across the cloud.
EDIT: using completely synchronous I/O (mentioned below) is a very clever solution, but it requires a process to know its inputs ahead of time. This may also cause cluster scalability issues, as now each "round" of inputs is gated by the slowest of the source processes.
If the framework you're using requires all I/O to be synchronous, and there's no way for a program to tell time or tell when an action would cause a delay, then there's no way for nondeterminism to develop based on timing.
I don't have any idea if ZeroVM is like this, and a framework that only allows synchronous I/O would have its own problems (basically you'd have to worry a lot about deadlock).
EDIT: To expand on this, you might still be able to do a lot even if cyclic interprocess data flows are forbidden. This is particularly true of database style applications, which are where ZeroVM originated.
[1] http://math-atlas.sourceforge.net/ [2] http://www.openblas.net/ [3] http://software.intel.com/en-us/intel-mkl
I work a lot with smallish data, where queries/processing may take 20 minutes or a few hours. Often I can split this work up significantly, but when most places charge by the hour it's rarely worth it. PiCloud are excellent in this area, and Manta looks interesting (but possibly a bit more expensive) but I'm not aware of many others. I love the idea of being able to start, with little overhead, a large number of short lived jobs. Particularly if I can run them locally or my own cluster.
I look forward to seeing more on this.
Lots of the recent activity is more about packaging and usability improvements, rather than theoretical improvements.
A quick overview: CoreOS is a super-minimal Linux distribution designed to be used as a base for applications. It's essentially equivalent to the JEOS buzzword from five years ago. It would run inside of Xen or KVM or VMWare.
Xen, KVM, VMWare, Virtualbox, you probably already know about- they provide a virtual machine, the operating system running inside of it (theoretically) can't tell it's not on its own hardware. Xen uses a 'hypervisor', which is essentially a very tiny, custom kernel. KVM uses the Linux kernel as the hypervisor, which makes a lot of sense- you don't have to reimplement all the years of hardware support and scheduling work they've done. VMWare and Virtualbox run as applications on whatever OS you provide. You lose out on some opportunities for clever performance hacks this way, but there are other advantages. VMWare ESX[i] is more like Xen & KVM, but I don't really know that much about it.
Containers (BSD jails, Solaris zones, LXC, and of course, HN's lovechild Docker) let you provide VM-like isolation and resource management between processes or groups of processes, but you only run one kernel. This means much less duplicated effort and memory, and Docker's AUFS lets you deduplicate your storage too. There are slightly more security concerns about this approach than full VMs, the Linux kernel (and others, but let's be honest about the target audience) has a long and ugly history of local privilege escalations.
ZeroVM is based on Google Chrome's NaCL, and that's about all I know about it. I would expect VM-like security (it validates machine code), and an environment that requires serious porting from POSIX. That said, if you use Python, Ruby, Mono-compatible .NET, or Go, the heavy lifting has already been done for you.
ESXi, on the other hand, was written to do away with the Console OS entirely, but it still has a fairly rudimentary shell. Many of the utilities are based on busybox and the idea is it should be stripped down with only really minimal functionality. It also sported something called the Direct Connect UI (DCUI) which is a curses based interface for doing things like settings up an admin password, reviewing logs and changing security settings.
Hypervisor is the wrong word- that's what you'd call Xen, VMWare ESX or the host KVM kernel.
Not really. Go used to have a nacl port, but that was years ago. It was abandoned when the nacl people decided to use a different method for isolating code.
Porting Go to use the new method would require writing another compiler, like {5,6,8,}{c,g}.
Determinism, OTOH, sounds interesting at least on paper. Is there any experience from tests with real applications in real world scenarios?
LXC starts as a general purpose Linux container with everything built in, and adds more isolation as development continues
ZeroVM starts with no general purpose Linux, and will add support as development continues
ie, LXC will work with what you have now, ZeroVM will eventually work with what you have now, but shims will have to be developed for everything, either in your code, or in ZeroVm's
IMO the future endpoint will have similar functionalty in both projects, but LXC will see more testing and use /now/
But, when you say tantalizing things like 'erlang-on-c', you raise the question: what does the clustering control plane look like?
One of the great things about erlang is that the cluster's got supervisors that receive execution-level messages (e.g. 'EXIT') and can then take whatever action they feel like. Is that control plane level exposed to ordinary containers?
And the other great thing about erlang is that the messaging model is either synchronous if you care (with return receipts) or asynchronous if you don't (fire and forget) -- and that richness turns out to have a bunch of good use cases. What's the ZeroVM story there?
And the other great thing about erlang is being able to trace out messages, especially when your synchronous architecture just took a dump on the sheets and is staring at you belligerently. Does ZeroVM have introspection figured out yet?
It looks like interesting technology, but I need a more concrete example.
But imagine that you're writing a commenting system, and want to sanitise stuff received from a user. Sanitising data is error prone, so you isolate the code in a new zerovm. If someone finds a way to exploit anything in your sanitising code, they might be able to write broken sanitised HTML out, but they won't be able to e.g. send queries to your database, or write to your disk, because the zerovm simply doesn't have permission.
And imagine the web server spawning a new zerovm for every request, that only has permission to talk to the inbound network connection and pass messages between that and a Rails zerovm for that request. If there's an exploit in the HTTP parser, that vm could be exploited, but it'd die at the end of the request, and would have no permissions to talk to the database server or write to disk.
And imagine the Rails zerovm similarly being split into pieces: Request handling might be done in one; authentication might be done in one.
The lower the startup costs, the more you can afford to chop the app into pieces and the more you can leverage that to benefit in terms of security (by reducing the privileges of each individual component) and scalability (by allowing distribution of the VMs across CPUs and across servers)
IPC is super limited: https://github.com/zerovm/zerovm/blob/master/doc/api.txt
You get nothing but /dev/stdin, /dev/stdout, and /dev/stderr by default. You can optionally make other resources (network, files) available, through a similar api.
AFAIK the idea as is to have as light of a container as possible so you can afford to throw your app at the data instead of throwing data at the app.
This may sound vaguely similar to how Linux containers and the syscall interface work, but it involves orders of magnitude fewer LOC written from the outset with a robust security design in mind. Compare that to the thousands of LOC daily churn in the Linux kernel, often written by people who are usually too busy fighting with shitty hardware to care about how their driver ioctl might be accidentally exposed to UID 0 running in a container, and even if they notice, might not even care.
If I have to run a database, say postgresql, how do I run it? Inside the ZeroVM or outside? To run the DB I would need to give it file system access?
Now if there is a security hole in postgresql, how is it guaranteed that files other than DB files are never accessed?
If you sub-divide your app in separate zerovms, whether per-request, or split it up further into functional responsibilities, then you substantially reduce the attack surface by ensuring that an exploit against any one part of your application can only exploit the specific subsets of data it is allowed to work on.
You can do this without zerovm too, but the more you reduce the cost and difficulty of spawning a new vm or container, the more finely grained you can subdivide your application, and hence the fewer privileges each subset of your app will have.
I can find this technology useful only in areas where you want untrusted 3rd-party code to run without worrying about what it will do.
And about "3rd party code". If we have two developers, each works on a different module of the same application, isn't their code is "3rd party" to each other?
On a serious note: the primary disadvantage of glibc is that it's really, really hard to change (and build times are slow). While it's already here, sometimes you want to port it to a new platform or a new ABI, and the adventure begins.
And glibc hasn't been slow for a year or two, since they defenestrated Ulrich Drepper.