For anyone who's worked with it, does the fact that it's pure userspace have any negatives? Performance? What about access to system resources, opening sockets, etc.
For anyone who's worked with it, does the fact that it's pure userspace have any negatives? Performance? What about access to system resources, opening sockets, etc.
That said, user space file systems are slower than in-kernel file systems; at least, according to the data I've seen. But, I suspect ZeroVM will be used in ways that aren't quite the pathological case of a user-space filesystem under benchmark load.
As for access, Linux has a nice capabilities system to allow user space to access privileged resources as a non-privileged user and without being in ring 0 (if privileged access is even needed, which it likely won't be in the cases where this will be used, I would think). Again, this would still be through the public APIs, and not by acting like the kernel and hitting the hardware directly.
I don't think, or get the impression, that this would be used for full-system virtualization. It seems to be more targeted toward an AWS Lambda sort of usage pattern; a micro-VM that spins up to serve one API request, or to act as a long-running tiny daemon to do some housekeeping task. It looks like a fancy fork(), to me, rather than a competitor to Docker or LXC (and especially not KVM or Xen, which can run a whole Linux kernel in the VM).
Almost. With ZeroCloud (OpenStack Swift + ZeroVM + appropriate middleware), you should imagine Lambda+S3 in the same service! Your "function" executes much, much closer to the location of the data, it doesn't require "shipping" from the storage service to the compute service and back again. You can take in an object, perform a transform (e.g. text search, encryption, transcode, etc) and store the result as a new object.
If you think about it, it's kind of the future of large scale computing, immutable dataset + immutable compute, that can horizontally scale to huge numbers of nodes.
Is that something that exists in a production form today? What's an example of the use case? I'm having trouble visualizing this "Your "function" executes much, much closer to the location of the data, it doesn't require "shipping" from the storage service to the compute service and back again." That sounds like going back to a monolithic model where data and functions are tightly coupled, but I assume I'm visualizing it wrong, since that would be moving backward.
If compute is separated from storage, then all of that video data has to be streamed over the network from a storage node to a compute node before computation can even begin; the data is "shipped" to compute.
Presumably the function you want to execute is vastly smaller than the data. It would require much less time and bandwidth to run the function on the same node as the data it's accessing; no network overhead. Assuming you have an adequate balance between compute and storage, you get much lower latency access to the data.
Some downsides include - running arbitrary code on your storage node means trusting your users or having very good sandboxing - you now have to balance compute and storage on any given node