Arrakis: The Operating System Is the Control Plane
usenix.org
usenix.org
That's the way IBM mainframes have worked since IBM VM, 1972.
There's a lot to be said for this. For historical reasons, most microprocessor systems have an I/O architecture that's far too much like that of an 1970s minicomputer - device registers with meaning determined by the peripheral. Minicomputers had that because they couldn't afford enough transistors to do mainframe-type channels. That problem went away a long time ago, but the legacy architecture remains.
You need both channels with access control and peripherals designed for channels to make this work without OS intervention. Peripherals may have privileged functions - user programs may be restricted to an address range on a disk, or an IP and MAC address on a network controller.
The Intel network cards (all the 10g ones, some 1g) at least have IP and MAC filtering, and support at least 64 virtual network cards for each physical port.
It has taken longer for storage, but NVMe has SR-IOV, which means you can split out virtual drives without the OS having to check block ranges. Not widely available yet, although Google cloud now has support[1].
Imagine this type of stack being used in an embedded system. I've worked on embedded projects that achieved high throughput but in most cases there were FPGAs and DSPs doing a lot work to help. Userspace-to-Kernel context switch delays have always been a latency issue with any Embedded Linux system I've worked on.
Arrakis looks like one would be able to achieve high performance without the need for FPGA and or DSP (depending on the use case of course).
Side note: Cool, I noticed they're using lwip from Adam Dunkels. He's an amazing programmer.
See also Intel SGX (https://www.virusbtn.com/virusbulletin/archive/2014/01/vb201...), disaggregation platforms (seL4, Qubes, Genode) and userspace networking (Intel DPDK, https://01.org/packet-processing).
Intriguing.
I think one issue you'd run across when doing something like that it that Node is not an application in the sense that Redis is an application. It's mostly a JIT-compiler for applications that are written in Javascript. Porting a Node application to run on Arrakis would probably involve changes to both Node and the Javascript application code, particularly if you want to run close to the hardware performance limits. That's doable, but I doubt you could preserve the APIs that Node currently exposes to applications. (The FS module would be especially problematic!)
What could be useful is to ditch Node's existing IO APIs and provide bindings to the Arrakis IO library directly.
Node is not a data specific application and won't live entirely in the data-plane. It is general purpose and libuv especially has a lot of moving parts that might blur between the data-plane and control-plane. I would imagine you would have better luck building a unikernel dedicated for node than to try and shoehorn node into Arrakis.
The bookkeeping associated with zero-copy often exceeds the copying cost. This was the curse of the original message-passing Mach implementation, which gave microkernels a bad name.
I'd like to see marshalling as a language feature. It's compilable, done often, and has an effect on performance. Many marshalling systems, from OpenRPC to protocol buffers, use a precompiler. But that adds another level of language.
Errr, that would be a major bug in capnproto. While it's definitely possible the software has bugs, it's certainly a design constraint that the sender absolutely cannot crash the receiver.
http://kentonv.github.io/capnproto/faq.html#arent-messages-t...
Can the kernel realize it and, if the hardware can manage isolation with OS semantics, step aside?
This isolation support in hardware reminds me of the "Software on Silicon" thing Oracle has shown on their new SPARC. Is offloading more and more application and OS level logic to hardware going to explode?