Kernel-Bypass Networking
godaddy.com
godaddy.com
Update: If you are like me and can't see the article at your localised GoDaddy.com website, you can (hopefully) select United States at the bottom of the page to force GoDaddy to serve you the US site.
As soon as you get the hardware to handle TCP reassembly and just wake the kernel up once per few megabytes of data sent/received, things scale well again.
There's work to do though - there are no systems around today that I'm aware of which can send data from SSD to a TCP socket (common use case for cache server) without the data itself going through the CPU (despite most chipsets allowing the network card to be sent data directly from a PCIE-connected SSD).
Not everything in this world is TCP. TFA mentions DNS, for instance (which as of this time is still mostly UDP based).
> There's work to do though - there are no systems around today that I'm aware of which can send data from SSD to a TCP socket (common use case for cache server) without the data itself going through the CPU (despite most chipsets allowing the network card to be sent data directly from a PCIE-connected SSD).
They certainly exist, they are just not available or known to the general public (hint: look for hyperscalers that manufacture custom network cards).
For NICs that have memory-mapped buffers, it could be just one transfer.
What am I missing?
Oh ? I’m quite disappointed, I thought that was the whole point of sendfile() ! What does it do if not that ?
sendfile allows for: SSD -> kernel -> NIC
which is already a major improvement
Where the data to be sent is non-contiguous in storage, the OS would probably have to make two requests (although those two requests could still be made simultaneously, so not requiring an extra wake-up/interrupt)
Not a generic solution, because it increases latency.
Which you could avoid by doing everything in user or kernel space.
On the other hand the kernel isn't standing still, overhead reductions have been trickling in over decades. sendfile, epoll, recv-/sendmmsg, all the multi-queue stuff, kTLS with hardware offload, io_uring, p2p-dma. The C10k problem was tackled in 1999, userspace APIs can get you much further today.
This is all not to mention locks, or that there are competing functions running in most distros (turn off irqbalance completely and watch your forwarding rate increase).
The low hanging fruit seems to have been picked as well - NAPI polling, interrupt coalescing, RSS + multique NICs + SMP, etc, are already out there, and we're still struggling to do 10G line rate in the Kernel...and data centers are moving quickly to 25/100G.
[1] Edited for terrible math - 10Gbps at line rate is 67ns per packet, 100Gbps is 6.7ns
From the kernel docs: https://www.kernel.org/doc/Documentation/networking/scaling....
""" While RPS steers packets solely based on hash, and thus generally provides good load distribution, it does not take into account application locality. This is accomplished by Receive Flow Steering (RFS). The goal of RFS is to increase datacache hitrate by steering kernel processing of packets to the CPU where the application thread consuming the packet is running. RFS relies on the same RPS mechanisms to enqueue packets onto the backlog of another CPU and to wake up that CPU. """
I have zero problem doing 25G and 40G ethernet at the Linux kernel in RHEL7 (I run Kubernetes clusters on them in fact). Non-ethernet (Infiniband) 100G+ line rate is also totally doable, but IB is an entirely different beast. I agree with you it isn't going to work in the long term, but it should be fine in the medium term. The long term Linux change is likely to just rewrite more and more of the existing filtering code ontop of the eBPF vm, which is fast as holy hell and is replacing large swaths of existing filtering mechanisms for a reason.
25G line rate is 37MPPS - are you saying you have zero problem forwarding at that line rate? I'd be very surprised if that's being done in the kernel. I'd be more surprised if you said that you're consuming bytes from the network at that speed with user space apps and no kernel bypass.
XDP (the forwarding approach built on top of eBPF) is limited to ~20MPPS, as well: https://www.netronome.com/blog/bpf-ebpf-xdp-and-bpfilter-wha...
Source: I operate a CDN with thousands of 100Gbe nics with a stock upstream LTS kernel, and minimal kernel tuning.
I just tested this with two hosts with 4.14.127 upstream kernel and upstream mlx5 driver, and mellanox connectx-5 card. Using 16 iperf threads
[SUM] 0.0-10.0 sec 85.1 Gbits/sec
That's pretty close with no tuning, and well beyond 10gb/s we mentioned earlier
https://events19.linuxfoundation.org/wp-content/uploads/2017...
https://www.redhat.com/en/blog/pushing-limits-kernel-network...
https://kernel-recipes.org/en/2014/ndiv-a-low-overhead-netwo...
"achieve 10 Gbps line rate at 60B frames"
"reaching line rate on all packet sizes"
Line rate is just bits per second. You have to add in a qualifier about packet size before you're talking about packets per second.
https://blog.ipspace.net/2009/03/line-rate-and-bit-rate.html
https://www.reddit.com/r/networking/comments/4tk2to/bandwidt...
https://www.fmad.io/blog-what-is-10g-line-rate.html
Also, for gigabit networks, ethernet packets are padded to at least 512 bytes because of a bigger slot size: https://www.cse.wustl.edu/~jain/cis788-97/ftp/gigabit_ethern...
64B is the minimum frame size in Ethernet, including interframe gap and preamble its 84B on the wire. It is the same with Ethernet, Gigabit Ethernet and even 100Gbit Ethernet, that source is not correct.
https://kb.juniper.net/InfoCenter/index?page=content&id=KB14...
Network hardware always quote PPS using the smallest sizes. And this makes sense for things like route and switch processors. Perhaps you are confusing that.
You should reread your link a little more carefully. From your link:
">However it is also important to make sure that the device has the capacity or the ability to switch/route as many packets as required to achieve wire rate performance."
The key phrase there is "as required." Almost nobody needs to sustain forwarding Ethernet frames with empty TCP segments or empty UDP datagrams in them. In fact many vendors will spec for an average size. Since packet size x PPS will give you your throughput, if the average packet size is larger you need much less PPS to achieve line rate.
These days line-rate just means sending enough traffic to fill the link at whatever rate you want. That's generally good enough for server folk since they just want to get you the cat pics ASAP.
That's not good enough for people running transit networks, since they care more about packets per second performance. Sending huge amounts of data is easy for them; what they really care about is PPS.
Aside, the next generations of router NPU's are trash in terms of PPS performance. I take that back, they're not trash. They're the trash in the dumpsterfire. That's how bad they are. We're fairly screwed there.
My guess is GoDaddy was looking at increased PPS performance either for DNS or maybe building their own DDoS mitigation framework (Arbor gear is pricey).
"For example, because one of the newest Cisco routers, the Cisco ASR 1000 Series Router, is capable of forwarding packets at up to 16 Mp/s with services enabled, it can support the processing of the equivalent of 10 Gb/s of traffic at line rate, with services, even for small packets."[1]
[1] https://tools.cisco.com/security/center/resources/network_pe...
The point of discussion here is that the Linux kernel struggles to do line rate 10Gbps. This was misinterpreted as "the Linux kernel struggles to do 10Gbps".
1538 bytes is 12,304 bits. 10,000,000,000 bits/sec / 12,304 bits/packet is 812,744 packets per second.
Now try it with 64 byte packets, which are 84 bytes on the wire.
14,880,952 packets per second.
And this is 10gbps.
Another example, you can saturate 100Gbps with just 4 iPerf3 processes.
The typical usecase are virtual network functions: think virtual switches/routers used to interconnect VMs or containers, or containerized VPN gateways etc. It is also used for high-performance L3-L4 load-balancers etc.
As pointed out by others, what is hard is to move small packets. TCP with iperf is not relevant for this kind of workloads. It is easy to max out 100GbE with 1500-bytes packets, but with 200-bytes packets not so much. This is why they communicate about PPS, not bandwidth.
There results seems low but it is hard to tell without knowing the platform or configuration. VPP can sustain 20+Mpps / core (2 hyperthreads) on Skylake @2.5GHz (no turboboost).
Thank you!!
Here is a relatively recent paper of the kind of work being done in this area,
https://iopscience.iop.org/article/10.1088/1748-0221/8/12/C1...
DPDK does have "kernel interfaces", so you can direct packets to the kernel.
There is also the additional overhead of VFIO/IOMMU's which can easily eat double digit perf vs doing the processing in a trusted kernel context.
So with a normal machine the cores/etc are balanced between processes, with DPDK/etc thats hard as some of the cores will be 100% consumed in those spinloops even if your experiencing 5% line rate.
TCP is genius for a WAN, but unlike most things designed in the past 25 or so years, robustness precedes performance.