Why do we use the Linux kernel's TCP stack? (2016)
jvns.ca
jvns.ca
I'd recommend understanding the stack before choosing to replace it (the old programmer's advice about understanding a gate before removing it). I summarized the TCP send path in my next book (sysperf2, with help from Linux network engineers); it involves:
- Socket send buffers
- Pacing
- TSO
- Congestion controls
- Nagle
- TSQ
- qdisc (optional)
- GSQ
- BQL
- TSO (NIC)
That's just the TCP send path.I'm happy to expand all these acronyms, but if any are new to you, that's kind-of my point. These parts have a purpose, and starting from scratch may lead you to eventually re-invent them all as you encounter the problems they solve.
And that poses a problem: to reach maximum performance requires many engineering hours to learn and configure all of these. You'd need a large gain to justify doing this: E.g., for a system doing over 1M packets/s.
This complexity may change. Look at BPF for observability: we built BCC and bpftrace so that using it was trivial. What trivial front-end/technology could we build upon XDP and io_uring to make deployment easy?
Can say why that is? I haven't really kept up with DPDK but I know there was lots of interest in it a few years back. What has changed?
Would using BPF to instrument the ioring be the way to go?
But if you invent them in a different order, you might just end up with a better system ...
Just look at some databases that were created by people who know nothing about databases when they started.
I bought a copy of BPF performance tools to learn more about the BPF interface, and it's really useful but quite focused on performance! I wish there was a resource of similar depth and breath about eBPF for classification/XDP.
It sounds like there should be a deep book on BPF internals (written by maintainers Alexei & Daniel), but it would be a tough sell: you really do need the maintainers to author it or be heavily involved to do it well, and it's a lot of work for a changing target, and without a large audience (lots of people will use BPF, but few will code it at that level. I'd guess 100 world-wide. Not a great case for a book!) Compare it to my BPF perf tools book: big audience (end users: developers, SREs, perf engineers, etc), and has a long shelf life (the tools are done, and work, and should keep working).
Pacing controls when to send packets, spreading out transmissions to avoid bursts that can hurt performance.
But is that really such a bad thing? Right now your choices are take it or leave it, or change operating systems. What if we could instead choose libraries suited to purpose, written by experts, like we do with encryption or compression or allocation or any number of common tasks?
You could then have a default implementation suitable for 80% of use cases, and specialized implementations to handle more complex requirements. You could even have per-NIC exclusive device access for a single application!
Are FreeBSD no longer getting any of those? It seems even Network Stack talks from Netflix are now Linux Focused.
It also makes me wonder..... does any of these matter once BPF takes over, quite literally.
The #1 reason I've heard is features. When you use the kernel stack you have a dizzying array of protocols, congestion-control algorithms, and other features available to you. You have all sorts filtering and rerouting and rate-control features, from old iptables to netfilter and eBPF. There's all sorts of familiar monitoring. User-space stacks never duplicate all of this, nor should they. They should support their own use cases, no more, but that means there will still be a lot of people who are better off with the kernel stack.
The #2 reason I've heard is that the kernel stack is known to work in a broad variety of environments. User-space stacks are often used in constrained environments where the device only has to communicate in certain ways with certain peers across a certain network topology. The kernel stack is known to work in a vast array of combinations simultaneously with less need for local debugging. There are also security aspects here. For example, a good TCP/IP stack will have features to mitigate DoS attacks both inbound and outbound, whereas some user-space stacks do things like "forget to implement" proper congestion control so they can DoS others. Whatever vulnerabilities do exist in the kernel stack tend to become known and fixed rather quickly. In your user-space stack? Good luck.
I'm not actually arguing that more people should use the kernel stack. There are plenty of good reasons to go the other route when appropriate. I'm just trying to present arguments I've seen and which should be considered when making that choice. The OP seems to look at things from a bit of a "single machine, single purpose" POV, which naturally skews the outcome toward user-space stacks, but in an era of virtualization and containers we need to consider multi-purpose machines as well.
The problem with those embedded solutions is often not lwip but that the integration and requirements between lwip, the OS and it's task system (e.g. FreeRTOS), and the actual hardware is not properly understood. Which can e.g. lead to function calls being made from the wrong context - which could work in the beginning - might lead to weird errors later on.
For this reason I would also recommend people who want to build connected devices to build on top of a more mature and streamlined stack - like the Linux TCP/IP stack plus associated drivers.
I also find the monitoring differences are important. If you have a userspace stack then things like lsof or packet capturing will work differently. If your stack is polling rather than getting interrupts you likely won’t even be able to look at cpu usage or load averages (each polling loop will run at 100% cpu looking for new network activity)
https://github.com/Mellanox/libvma
https://blog.cloudflare.com/a-tour-inside-cloudflares-g9-ser...
https://blog.cloudflare.com/l4drop-xdp-ebpf-based-ddos-mitig...
1: Snap: a microkernel approach to host networking https://blog.acolyer.org/2019/11/11/snap-networking/
At the most basic level you can run an existing binary under Onload (which basically uses an LD_LIBRARY_PATH to direct the socket system calls to their own stack) and get a significant performance improvement, especially if you are careful about how the card is configured, and to ensure your app is running on cores on the same socket that the network card is connected to.
Taking it further, you can use a zero copy API to get inbound packets whilst they are still in the network card buffer, and do very cool things with both hardware timestamping and pre-caching of packets on the network card memory (sort of like a message template) which can reduce the latency to sending responses if you have a set of stock responses you want to supply.
Anyhow, do check them out if you are into that sort of thing.
You can get experience with ordinary Intel hardware and AF_XDP or DPDK.
Might also want to look at Packet. They offer bare metal cloud servers by the hour, some of which have Mellanox NICs in them[0]. I'm less familiar with Mellanox than SolarFlare, but I think they have a similar feature set.
My employer's customers don't use Mellanox anymore, so I haven't. I don't know why they don't, but would not be surprised if price had something to do with it.
I say all this because I highly recommend playing with Linux’s raw packet interface and learning the structure of IP packets to do so. You can really gleam a ton of interesting info about it that way. Writing your own TCP stack was outside of what I needed but even the comparatively much simpler task of routing UDP packets was really fun and rewarding.
It sounded like the latter but for anyone unaware such sockets require root/special privileges. It’s cool but of more limited general value and can be a security exploit as a user space exploit can launch really nasty network attacks that are otherwise impossible (hence why it’s not made available to user space).
... if you don't update your kernel.
Rolling a custom TCP stack is only interesting when the amount of work per TCP connection is minimal and you're in the territory of handling half a million or more concurrent connections. Think of caching servers etc.
However if you had that many concurrent connections to some website you'd already be comfortably sitting in the Alexa top 1k of internet services.
If I needed to do load balancing via an intermediate server, it would make more sense for me to load-balance on the raw (i.e. TLS traffic) TCP layer.
In my case, I am getting UDP (multicast) packets at a sometimes ferocious rate, that need various kinds of processing done, and all captured, with nanosecond timestamps, to disk.
The key is to get the kernel not involved at all. You use NICs from Solarflare (now a Xilinx company, which they used to be a customer of); or Exablaze (now owned by Cisco); Netronome; used to be, Mellanox (now owned by NVidia); or even Napatech (expensive). The NIC driver sets up a ring buffer in DMA memory and just starts dumping packets into it in real time. Each packet gets a little bit of metadata: a nanosecond timestamp, byte count, checksum. The NIC might be filtering by IP address or port, to distribute incoming packets to different ring buffers. The ring buffer is typically a few megabytes, enough for packets to be there for a few ms before they get overwritten.
Your program has to watch this ring buffer for updates, and do whatever it needs to do before the packet gets overwritten. I memcpy them to a big-ass ring buffer, say 8GB of hugepages, and then other unprivileged processes can pick over them for interesting bits, with more leeway for stalls, and can be started and stopped independently.
The process watching the NIC has to be protected against interruptions from the kernel, which involves a mess of kernel boot options -- nohz_full, isolcpus, rcu_nocbs -- because kernels are very jealous of their privilege to stall any process and steal its core for their own purposes. The program needs to do no system calls after startup, and not write to any mmapped memory backed by actual disks (/dev/shm and /dev/hugepages are ok), or the kernel will stall it anyway, boot flags notwithstanding.
Typically each NIC maker has its own kernel-bypass driver and library, often open-source, that understands its ring buffer. Usually they provide a .so your program can LD_PRELOAD to divert regular socket calls into their library, that you will ignore unless you want to, e.g., send out TCP traffic.
NICs have unique features. Intel and others implement a more-or-less portable DPDK interface to their library. Solarflare provides an Onload implementation, and pretty smart hardware filters. ExaNIC has less-smart filters, but delivers packets 120 bytes at a time, so you can start work on a packet before it has all arrived. A Netronome NIC can run eBPF code on a core in the NIC, against packets not even copied to host memory yet. Napatech lets you mess with the filters from the command line, while it's running, and can send packets on a nanosecond-resolution schedule.
Most let you queue up packets and trigger sending certain ones on a dime. You could have a dozen packets with different possible choices, and send only the one you later determine is right.
Lately there is a kernel service, AF_XDP, that is supposed to be a portable, kernel-maintainer approved way to do some of what the proprietary libraries do. I haven't tried it.
Getting reliable nanosecond-resolution timestamps is tricky. Nowadays everything is referenced to atomic clocks on GPS satellites. So, you need a GPS receiver, and a way for the NIC to know what it says. ExaNICs have a receiver on board. Often there is a connector for "PPS" input, expecting a clock rising or falling edge at a known offset from the second boundary. A protocol, PTP, provides ~microsecond resolution, but burns one of your 10Gbps ports. Some switches will process PTP, PPS, or GPS, and tag packets with various non-standard annotations.
If you need to do trickier things, several NICs have FPGAs you can program yourself.
I'm not sure what level of timestamp precision is required for those. The solutions also likely don't require 10Gbit/s.
Arguably, it amounts to automated stealing, but is 100% legal.
All the NICs have FPGAs on them nowadays, to do regular IP-related stuff, so it is just a matter of reprogramming them.
In practice, microseconds is better than enough, and all you get from PTP anyway.
Why not? Because you don't trust your skillset to write one? Because you don't find time to write a good enough one? Because the existing ones are good for you? Those can all be valid reasons.
But I don't buy into the general "software you shouldn't write yourself" thing. If everyone would follow that advice, we would have a lot less progression and improvements in the field.
I have written custom HTTP implementations for all protocol versions and would be find deploying them in production - that's what they had been written for. I would also be fine writing a custom TCP stack or TLS stack if something requires that.
But before doing that I would need to make sure I've done all my due diligence: Have I fully understood the problem - and is my solution solving it? Do I understand all the edge-cases, which might lead to serious availability or security issues later on? Is my code well tested? Obviously there is always something that one only learns later, so those checks will never be exhaustive. One good minimum check is reviewing your own implementation and tests against the one of existing implemetations, and see where it improves something and where it might miss out on something.
You will also need a goal, on where you want to improve over the existing solution. If you don't have that goal, and have no way to measure/quantify it, then don't do it. Also think about how much that goal really matters for your end product. If it doesn't matter at all - then either don't build the thing or build it as a side project or learning project.
Linux Kernel TCP stack is not just a faithful implementation of the spec. It's also all that hard learned experience about Undocumented Real Internet that is constantly evolving and what must be dealt with.
A bug in TCP stack can cause lots of problems for users. Taking responsibility of that is also huge task.
https://www.saminiir.com/lets-code-tcp-ip-stack-1-ethernet-a...
OS's are doing too much in my opinion and the problem is that the status-quo of applications expect them to do it, so theres no clear way out of this.
If im not mistaken, theres a paper with the 'exokernel' design from the late nineties, built from a NetBSD, that describes a little bit about this approach, if you are curious.
Our OS's should be like that, and i bet they would be so much better giving they would be able to concentrate in being good in a much smaller spectrum, while applications would probably negotiate less with the kernel, giving us a better performance and stability overall.
TPC and sockets would be a library, and once we need to prototype new shiny stuff, it would be much better to get the needed adoption by just linking to the appropriate library.
Projects like QUIC and gRPC would be much more viral, and i bet they could even appear earlier if some part of the stack were not frozen right on the kernel.
I guess, this will become more evident with time, as a lot of those technologies on the kernel will become deprecated, while in the heat of modernism it look a great idea to stuck that thing that "everybody uses" on kernel space.
Also, one of the user-space stacks seems to be available via binary blob only.
I'm hoping io_uring will be good enough for higher QPS.
One concern with custom HTTP handling is that HTTP/2 and HTTP/3 seem to be non-trivial to hand-roll. HTTP/1 is a pretty simple request-reply design.
Sure, the kernel is implementing the TCP/IP stack, but once the right socket has been identified based on IP address and TCP port, and once the packet has been re-assembled, all checksums verified, and the ACK sent to the sender, the bytes of the TCP payload are delivered to the user-space read() or equivalent call, and it's up to user space code to interpret that as an encrypted HTTP request or response, decide if the entire request has been received etc.
This is why OpenSSL or GnuTLS are libraries that you can find on your system. And the actual HTTP implementation is part of Apache or NGINX or other web servers, or by browsers and other clients.
Edit: maybe I still wasn't clear enough. All I'm saying is that, whether the TCP/IP stack is in kernel or user space, the specific HTTP(S) handling is always in user-space in all OSS that I know of - so there is no choice to be made for HTTP itself in this regard. You can put any HTTP 1, 2 or 3 implemention that you like over either DPDK or the traditional Linux TCP/IP stack.
HTTP/1 has enough edge-cases to balance that scale back out:
* chunked transfer encoding; there can be extensions with chunks
* trailers (headers that appear after the body)
* Expect & 100-continue can mean >1 "response" for a single request; also, you can't process this header for a HTTP/1.0 client, as it didn't exist then. (And you can't send it inadvertently to a /1.0 server, as you'll never get a 100-continue back!)
* HEAD is an absolute anomaly in terms of handling the response-body; you have to know you sent the HEAD, as the response will likely contain something like "Content-Length: 100" but there won't be 100 bytes of body to follow.
* If you handle obs-folded headers, obs-folded headers. Though, they're now an optional part of the specification, as they're obsolete. (But in that case, one still has to correctly 400 the request.)
Some of these exist, of course, in some manner in HTTP/2, but they're often more straight-forward to parse due to the binary nature of the protocol.
Curious, why would it need a second VPC? DPDK, albeit on AWS, allows me to attach a second NIC to the instance on the same VPC.
From what I understand, complete kernel bypass networking is only really feasible if you only have a single process that is using the network. And with a monolithic kernel, the TCP stack needs the kernel for port allocation and load balancing the demands on the network interface.
Okay, that makes sense, but why can't most of the processing be offloaded to userspace, with the kernel just doing those specific tasks? Port allocation should happen very rarely, and throttling should only ever happen when the network interface is saturated. Wouldn't it be better to limit the kernel to just those tasks, and offload the rest to userspace?
------ [0] https://www.cs.rice.edu/~alc/old/comp520/papers/U-Net%20A%20...
The TCP/IP stack will interpret the Ethernet headers, drop the frame if it is not addressed to any of the machines, check checksums, handle ARP requests etc. Also, the stack would buffer the frames until it can reassemble an entire IPv4 datagram out of them. Then, it will perform similar checks on the datagram destination address as above.
Finally, with the IP datagram in hand, we can decide whether to hand the payload over to the TCP handler, the UDP handler, the ICMP handler, or perhaps others. Of course, we need some TCP knowledge as well, to allow multiple TCP handlers to co-exist, choosing the right one to forward the packet to by port.
At all of these steps there are other concerns that need to be handled. Network security tools should get to look at the IP addresses and MACs to decide if the packet is accepted at all. Raw sockets created by user space can receive the Ethernet frames or IP datagrams directly, before other processing. There can also be contention on those as well (e.g. Maybe 2 processes both open a raw socket on the same NIC - who gets the bytes?).
Multiple NICs on the same machine also need to be shared and load-balanced. The Linux kernel for example handles routing between these, even applying networking decisions such as advertising IPs from one NIC on other NICs when it thinks it can improve delivery rates.
Now, all of these can and have been done in user-space for various reasons. But it's important to understand that there is a large amount of processing. Also, given the importance of networking to IPC and security, there are also security considerations at almost every step of the way. For example, if two privileged processes on the same machine are communicating over a virtual NIC with static IP allocation, can they expect their communication to be isolated from a non-privileged process running on the same machine, or do they need to use an encrypted&authenticated scheme? A centralized TCP/IP stack can make that guarantee, a per-process stack can't.
That’s a huge part of why the kernel implements it. The other part is that network stacks typically can be more tightly coupled with network drivers for performance (again trusted context so less/no validation needed for the buffers being shared and whatnot). Now there are obviously cases where user space can do a better job of buffer management for better latency/throughout so the perf aspect isn’t that conclusive.
If you don't care about the order you are receiving things then you want UDP.
Is it better for userspace to do the sequencing? Maybe given NUMA - a process might be "closer" to its memory than the kernel and therefore could eat its traffic in the correct order locally faster than being fed by the kernel.
In short: You want to have something that does the multiplexing and demultiplexing of packet streams that belong to different applications. E.g. all packets arrive from a common NIC, but depending on the target port they have to a different application.
And on the sending side you want to have some fairness, and common congestion control. If all applications could write directly to the TX path of the NIC it could easily be that packets from one particular application which has less CPU time would commonly get dropped.
In addition for that having a common place like the Kernel allows for some nice monitoring/instrumentation purposes. And it makes sure peers get a reasonable response (even if that's only a RST) if the application is not started.
The difference between non-DMA and DMA is staggering.
In a large deployment, anything that doesn't fit into the common update/remediation workflow is going to require special accommodation in code. Is the engineering cost plus the hardware cost worth it to the customer? Sometimes still yes, but more often no. There are many examples of companies who found out the hard way that the market for this kind of thing isn't big enough to recoup their own development and other costs.
P.S. This is very similar to the arguments for/against hardware RAID controllers. For whatever reasons, rightly or wrongly, those are also steadily losing popularity. Software really is eating the world.
P.P.S. In some cases, e.g. Amazon, the "smart NIC" approach is the common workflow, so the color of this argument changes. OTOH, it's also worth noting that the kinds of network filtering/virtualization/whatever that Amazon does is very specific to them and has nothing to do with any standard. They dedicate staff to support it. Bespoke ASIC/FPGA approaches aren't the same as a market in which you can sell them.
I was under the impression that these "SmartNICs" were already commonplace in prominent public clouds. Is this a completely unrelated use-case?
Linux already supports offloading even arbitrary eBPF/XDP programs to compatible NICs, which is already extensively used for DDoS mitigation (Cloudflare) or even Load Balancing (Facebook).
I'd be surprised if it didn't leverage these mechanisms to offload the kernel's own network stack as well, but haven't really checked to confirm...
However, specialized stacks for certain purposes are a relatively common thing. For example, most L23 network testing solutions use FPGAs to generate and receive massive amounts of (almost) stateless traffic (think line rate for 400GE using just 2 machines), while performing certain kinds of analysis on it (latency, loss). But these are usually just Ethernet frames or maybe IP datagrams with random payloads.
Lots of userspace TCP-stack-equivalents will be running soon, though they'll probably consists of just various versions of a couple of codebases -- your average webserver won't implement it's own low-level stack.
Did you mean to say thrashy (not that it's wasn't also trashy)? Because IIRC that was the main problem. There were issues with, for example, the memory mapped area leaving the caches of the resident core that the process lived on; and also generally BeOS was terrible at pinning processes to CPU cores. And I recall that the sentiment was also that the network stack needed low latency in general and that it was rather difficult to get a good implementation of a userland network for these reasons.
I imagine that at least some of these concerns may play out in more contemporary userland network stacks. Possibly, worse so if you don't know what is going on with noisy neighbors in containerized or even vm-ized deployments (though I would guess that Amazon/GCP but maybe not Azure does a good job of quotaing/throttling vms)
Maybe that sounded open minded and "oh shiny new feature" in 2016, these days it sounds like embrace and extend...
If instead you want a networking stack written in Python, for WSGI based frameworks like Flask it is enough to write a WSGI server that instead of using the `socket` module uses your custom stack. It will be compatible with the vast majority of Python web apps thanks to WSGI.
That being said it's very unlikely that you will notice any performance improvement. In the case of a networking stack written in Python it would even be way worse than the kernel implementation.
I'm not competent or eager enough to want to deal with, say, Path MTU discovery on my own
What makes the Linux kernel TCP stack expensive?
2. It hasn't even finished standarization
3. It can be up to 5x less efficient than using TCP+TLS (personal measurement). If you don't require TLS it can be even worse.
4. It doesn't even answer the question that the article asks: You still have lots of options on how to get packets for your Quic endpoint to the application. You can use regular socket APIs, you can do some trickery using parts of the kernel stack (using eBPF). Or you can even do even more kernel bypass and using AF_XDP to get data into the app.
Anyone know of a tool/command that can measure how much traffic (mb/s) is being sent to an application?
Related: trickle (per application bandwidth shaping) and wondershaper (per interface bandwidth shaping) are great for testing.
`nethogs` is also useful.
If you'd like to do something more specialized, you could build something with a couple of lines in `bpftrace`, but you'll have to put in more work there if you're not familiar with it.
The book "BPF Performance Tools" by Brendan Gregg is a good introduction there. Of course, for a simple overview, that would be way overkill, though :)
I whipped up a simple shell script for the polling thing.
Usage: ./traffic.sh <pid>
Copy-pastable for your terminal (dollar signs signify your prompt, obviously):
$ cat <<'EOF' > traffic.sh
#!/usr/bin/env bash
pid=$1
set -euo pipefail
while true; do
date +'Time: %H:%M:%S'
sed '/:/!d' /proc/$pid/net/dev
echo
sleep 1
done \
| awk '
BEGIN { MB = 1024*1024.0; }
/^Time/ || /^$/ { print $0; next; }
seen[$1] {
rx_mbps = ($2 - rx[$1]) / MB;
tx_mbps = ($11 - tx[$1]) / MB;
printf("%10s Rx: %8.1f MiB/s, Tx: %8.1f MiB/s\n",
$1, rx_mbps, tx_mbps);
}
{ seen[$1] = 1; rx[$1] = $2; tx[$1] = $11; }
'
EOF
$ chmod +x traffic.sh> TL;DR: systemd now can do per-service IP traffic accounting, as well as access control for IP address ranges.
It can be enabled using:
[Service]
IPAccounting=yes
systemd won't, AFAIK, give you per second values -- just the aggregate totals -- so you'd have to constantly retrieve the updated values and calculate the differences yourself. The totals can be obtained programmatically via the D-Bus APIs. They'll also be shown in the "systemctl status ..." output or you can grab just the value(s) you're interested in with "systemctl show ...".However...
systemd's functionality is basically just a thin wrapper around eBPF (in the kernel, since sometime around 4.8, I think?) so you can do it without systemd too.
In fact, the author of this article has also written a nice introduction to XDP and eBPF [1]! That article includes an example "XDP program" which counts the number of packets that have been blocked from a list of source IP addresses -- it could be modified fairly easily to do what you want, I think.
---
[0]: http://0pointer.net/blog/ip-accounting-and-access-lists-with...