Snabb Switch – A toolkit for solving novel problems in networking
github.com
github.com
Once you do some meaningful work (say, HTTP protocol decoding), this figure will be a lot lower.
Anyway I suppose the main target of this project is to help developers of packet switching and load balancing software.
If this kind of library can help us do that, that could save us a lot of time and effort. Time we could use to work on the next bottleneck :)
Snabb means fast/quick in Swedish.
(I imagine "Snabb" is closely related to "snappy" in English.)
Besides, dedicated hosting for a 1U or 2U rack machine is not expensive. If you have enough traffic that it's worth building your own TCP solution, you already have enough servers that you're spending a considerable amount of money on.
The general idea is to have hardware with direct virtualization support (which is increasingly available on commodity hardware), then have a 'control plane' of layered, virtualized syscall APIs that configure a 'data plane' of virtualization-aware hardware. Permitted I/O operations occur just as if they were on bare metal, with asymptotically zero performance overhead, because you can process an arbitrary amount of data without invoking any code at the OS/hypervisor/cloud provider layer.
For example, my rented, virtualization-aware CPU allows me to run any non-privileged code that stays within a certain block of address space; my rented, virtualization-aware NIC allows me to send and receive any ethernet frames that match certain header bits; and my rented, virtualization-aware disk allows me to read and write to a certain range of LBAs. The nth-layer OS or cloud host or whatever can come in and alter these permissions at will but it need not examine every single syscall to see if it conforms to policy.
The 10G interface is presented as 64 bits of data every clock cycle, so if you can express your algorithm entirely as a byte-by-byte state machine you can parse messages with minimal delay.
I know Intel has software for their NIC's to enable userpace networking, but I don't think any HFT uses it. It may be all right for hobbyist experimentation. I left the industry about a year ago, so my information is a bit out-of-date, but my former employer used either Solarflare cards with the OpenOnload stack (very good cards and awesome software) or the Mellanox CX series with VMA stack (amazing hardware with mediocre software. VMA is now open-source, so perhaps that situation has improved).
Note that these cards are a few thousand dollars each, so out-of-reach for hobbyists.
For the latter (which matters in routers, for instance), netmap runs on everything (either natively or with some emulation). The intel 10g cards (which DPDK of course supports) are around 3-400$ i think, the 1G cards start around a few 10's of dollars (not that you need userspace networking at 1G).
Also, just in case the author of Snabb Switch is reading this, it would be really helpful to provide a short example program of how to use this. A brief look over the documentation didn't show anything. (It's possibly I just missed it. In which case it would be great if you linked to it in the README or otherwise made it prominent).
https://github.com/SnabbCo/snabbswitch/tree/master/src/desig...
Here is the driver hack that makes it possible: https://github.com/SnabbCo/snabbswitch/blob/master/src/apps/...
http://info.iet.unipi.it/~luigi/netmap/
netmap / VALE is a framework for high speed packet I/O.
Implemented as a kernel module for FreeBSD and Linux enabling memory mapped access to network devices.
[1] Look for snabb switch here: https://cast.switch.ch/vod/channels/2i5k459xe3
There is also a podcast here http://blog.ipspace.net/2014/06/snabb-switch-and-nfv-on-open...
Because Snabb is specialized around a few specific NICs and virtual IO interfaces with specific features, it is able to do things like set up the network hardware to be memory mapped. As in, after some setup, Snabb can bang on the bits of a specific memory range and that will change bits in the network card state without the OS or even the driver being involved. This means you are not calling send() or select(), you are dealing directly with the Intel NIC hardware interface.
It sounds a lot more like game console engine programming than node.js programming. As a game console engine programmer, that's pretty interesting to me :)
"I'm the Snabb Switch originator.
The project is new: I and other open source contributors are currently under contract to build a Network Functions Virtualization platform for Deutsche Telekom's TeraStream project [1] [2]. This is called Snabb NFV [3] and it's going to be totally open source and integrated with OpenStack.
Currently we are doing a lot of virtualization work from the "outside" of the VM: implementing Intel VMDq hardware acceleration and providing zero-copy Virtio-net to the VMs. So the virtual machine will see normal Virtio-net but we will make that operate really fast.
Inside the VMs we can either access a hardware NIC directly (via IOMMU "PCI passthrough") or we can write a device driver for the Virtio-net device.
So, early days, first major product being built, and lots of potential both inside and outside VMs, lots of fantastic products to build with nobody yet building them :-)"
What are they talking about?
Given that servers are getting 10Gb, 40Gb or more, this is becoming more and more of an issue. The hardware is capable of it, but the abstractions are getting in the way.
Reading it again, it sounds more holistic -- that a simplified, monolithic application with a local, lightweight database is much higher performance than, for instance, an n-tier, distributed database type solution.
where userspace needs to look at each packet
That benchmark is like those naive "look how fast my web server is when it returns just status code 200" comments. Even if we accepted that the overhead was anywhere close to the linked, which it isn't, the moment you actually do something with the packets those savings disappear into rounding errors.
On passing, people typically use the term 'os-bypass' but netmap relies on the OS for protection, synchronization, memory management etc -- all things that the OS does well and i find no reason to reinvent.
As a test of max throughput, it is a horrible test. I don't have the motivation to prove it, but I would be surprised if more than 5% of the CPU load actually went towards networking, the rest time calculations and interval tests.