Transitioning into kernel space shouldn’t take that long so a lot of needless instructions must be getting executed.
Transitioning into kernel space shouldn’t take that long so a lot of needless instructions must be getting executed.
Ref: https://opennetworking.org
Not just the datacenters, but even 5G deployments, which would be exclusively IPv6 based, would sport user-space processing of packets and flows.
Ref: https://opencord.org
There are inherent problems with relying on the OS in an elastic environment. Kernel's TCP/IP stack could be made to expose very specific APIs to help with NetworkFunctionVirtualization (NFV) and Software-defined Networking (SDN), but the advantages are not clear:
1. Vendors to rely on Kernel devs to do the their biding for their NFV/SDN deployments, and wrestle with various other priorities Kernel might already have.
2. Build/deploy the updated Kernel fleet wide, with those changes (which now include other Kernel changes vendor may or may not want).
3. Cater to strict Kernel community guidelines and expect to be driven by the established processes governed by designated gatekeepers, costing vendors time and speed to market.
With a userspace solution, esp since it allows vendors to be performant by simply interfacing with NICs, the vendors have a measure of control on their own destiny:
1. Deploy at will, at pace.
2. Ability to do custom packet/flow (virtualization), on-the-fly (SDN) even, without requiring any assistance from any other part of the system not in their control.
3. Reduce external factors that might influence their timelines or priorities.
The generic problem, however, with the Kernel's TCP/IP stack needing improvement is orthogonal to this. That effort will have little to no bearing on the direction the Networking industry is moving, imo.
---
Google Maglev (network load balancer): https://ai.google/research/pubs/pub44824
Dropbox MagicPocket (bypassing OS' VirtualFileSystem for storage at scale): https://blogs.dropbox.com/tech/2016/05/inside-the-magic-pock...
Facebook Haystack (an example of problems working with VFS at scale and a solution to it): https://code.fb.com/core-data/needle-in-a-haystack-efficient...
NFV: https://www.electronics-notes.com/articles/connectivity/nfv-...
VNF: https://www.electronics-notes.com/articles/connectivity/nfv-...
read() interfaces expect the kernel to put the data where you asked. its hard but not impossible to coordinate with the NIC to demultiplex the packet and get it into the right place without a copy. almost all of the time you just want to see the data, you dont care that it lands in this particular spot.
getting control flow in and out of the kernel is expensive. target address need to be checked. registers often need to be saved.
things like the kernel firewall take a certain amount of work to determine the packet status
there are lots of mitigations, but at some point you run into a conflict between general purpose multi-process functions and performance.
if you ran your application on a dedicated arm next to the nic I bet you'd see a difference. maybe that shouldn't be so hard.
So a simple UDP recv() + copy from network card buffer to userspace should be about 3-4usec. In the real world, it's more like 6 or 7usec the last time I timed it.
Admittedly, the unix API could use a refresh. A lot of unix calls like getaddrinfo are perfectly happy to allocate and give you a buffer. A "fast_recv" could easily do the same and just hand you a page from a kernel page pool directly that you then have to free.
By avoiding the buffer copying, you could get something in the same ballpark as kernel bypass but without all the bespoke code and frameworks in every application. And in the end, that's kind of the point of having an "Operating System" in the first place.
Not sure where the Linux kernel lags (may be at scale it kind of tails out?)-- if you read Google's paper on their network load balancer, Maglev, where section 5.2.1 specifically calls out 30% performance degradation without bypass on smaller packet sizes.
> By avoiding the buffer copying, you could get something in the same ballpark as kernel bypass but without all the bespoke code and frameworks in every application.
May be this article from Cloudflare helps paint a proper picture of why certain kind of applications may need to keep bypassing the kernel forever because their requirements are very specific and not because the kernel can't be made to go faster (ex: a virtual router/switch, or a load-balancer): https://blog.cloudflare.com/why-we-use-the-linux-kernels-tcp...
(If you're Google, you hire the kernel hackers, but if you're trying to make money off a product and ship it next quarter, you buy the NIC)
It's a massive waste of resources. If the kernel could figure out DRI for graphics then they can figure out universal passthrough networking. All the bypass libraries seem to use the same buffer pool + spinlock slot architecture already.
And by your logic, how does anything ever get into the kernel? Nobody should ever contribute anything because "upstream".
That's lack of leadership for you... Kernel people should be pushing back more on such features
The linux eco-system does not begin and end with a RHEL/Ubuntu binary distribution built to run on every crappy PC in the last decade.
2. focus - kernel keeps accumulating abandonware code for gimmick features/drivers