Handling TCP in the kernel has some overhead due to system calls.
Also, the way sockets are designed does not make them very scalable, as you have lock contention on the TCP state machine. The SO_REUSEPORT feature introduced in Linux 3.9 solve some of these lock issues, but the kernel TCP stack is still not fully parallel [2].
--
[1] https://github.com/RaphaelJ/rusty
[2] https://raw.githubusercontent.com/RaphaelJ/rusty/master/doc/...
Yes, and this project, if moved into kernel-space, would entirely replace that stack and its state-machine.
> system calls
You’ve still got the overhead (context switches and memory copies) of getting the IP packet out of/into the kernel, which I don’t think is all that much less than the overhead of getting a TCP packet out of/into the kernel.
Really, what you want is SR-IOV to allow the user-space process to do direct Ethernet DMA to its own dedicated network card. No copies at all!
But if you’re willing to do that, then the application is basically acting as its own kernel... so why not just admit that, and instead of writing a user-space process that has half the features of a kernel, just either 1. write your logic as a Linux kernel driver, or 2. compile your program into a unikernel framework? Then your VM-nee-application’s host can be a proper VMM like Xen or ESXi, where it’s easier to configure that SR-IOV dedication as part of your VM-nee-application’s workload configuration.
For this reason, I’ve never understood people trying to do things like this “in user-space.” You’re playing at being a kernel—with all of the problems of being a kernel—without the ability to rely on an existing, well-written kernel as a basis for your logic (like e.g. the parts that handle the L1-L3 layers of the network stack, which you aren’t changing much.)
Because the kernel is still a lot of other things for you other than networking - I think it's a stretch to say that all user-space networking makes your work "half" of that of a kernel. And, you're not necessarily the one doing it. You may be an application, and your user-level TCP (including kernel bypass) may be a library from someone else. But to your general point of now you are now well past the city walls, and may run into trouble, I agree. I assume that this sort of thing is only done by a small number of people.
Microsoft calls this receive side scaling, and it's also available in FreeBSD. On FreeBSD, this really helps with tcp data packets, but connection setup still has bottlenecks; there's not an api for setting up outgoing sockets to align with the cpu you're on, but it's possible to do it with manual port assignment on outgoing connections.
Every mainstream OS today predates mainstream SMP. While support is built into each, it's just ... support. They just were not made in the world where every system is SMP. To see how big a design difference that makes, well, look at Erlang or Go.
I think the Haiku kernel was also? At any rate it certainly is now, in much the way DragonFlyBSD's was if nothing else.
So, was it not designed for concurrency and they just somehow fixed a lot of it later? That retrofits usually don't work is why I believed claims that concurrency was a design goal.
But ... I didn't say anything about BeOS in my comment? I was talking about Haiku, which has very different origins than BeOS, especially on the kernel front where the internal architecture was and is pretty different from the Be kernel.
I actually don't know how much the Be kernel was designed for SMP from the first days; I think it was but I'm not sure. At any rate it definitely did have better concurrency for desktop usage than anything else at the time, I believe.
[1] http://www.mellanox.com/page/products_dyn?product_family=209...
100GBit would usually put enormous strains on the CPU (to the point it’s pretty much only doing that), which is taken away entirely by using this.
As it turns out, hardware appears to be much faster at these types of things than general purpose CPUs. :)