Zero-copy network transmission with io_uring (2021)
lwn.net
lwn.net
So I dug into it .. and well I'll have my own library soon. I should be able to send UDP w/o congestion control sometime this week.
eRPC uses DPDK (100% user space NIC TX/RX control) plus the author's own other ideas to get performance. Since I'm getting into DPDK way, way late in the game, I hope and believe DPDK is and will be the better way to go than to turn to kernel fix-ups.
Getting the kernel out of the way with pinned threads seems -- or more exotic with special H/W if preferred --- is cleaner if one can develop from scratch. DPDK is all day long zero copy, and always has been.
This library will be a part of something bigger. For many frameworks, a key architecture point is:
I got a RPC/packet/message. Ok, now what?
* process it in-place e.g. on the thread that was doing RX?
* if I don't delegate am I making back-pressure?
* delegate it to another core? How to efficiently do that? Hint: https://blog.acolyer.org/2017/12/04/ffwd-delegation-is-much-...
* if I delegate ... to whom? Maybe I'm partitioning?
* If I delegate how do I get a response back?
In DPDK I believe these are easier to decide and well-manage in code.
>execute the code inside the NIC driver
this is ... a NIC+FPGA which can execute code?
I know of some issues myself that I plan to fix, e.g., using macros to decide which Transport class to use -- this should just use virtual methods; messy inter-mixing of congestion control and packet I/O; some "performance optimizations" like preallocated buffers that are just not worth it.
I do and will continue mention your paper (which I also found on HN) and your repo because it has a lot of really good ideas ... ideas which I'm forced to engage to mastery.
I haven't yet had time to focus on how you deal with packet loss via and with congestion management. Anything you can do --- I'll have to reread your paper (well written BTW) to pull my own weight here --- is a real benefit to the community.
I think the other questions are more ones of library design and less about the packet API one uses. On top of all kinds of APIs one could build singlethreaded state machines which deque packets and process them. Or other architectures if required.
- userspace networking which avoids copying in the kernel and expensive context switches (both advantages of io_uring) and some costs of sockets apis[1]
- specialised NICs that can e.g. do TLS in hardware
- specialised NICs with on-board FPGAs to which some custom processing may be offloaded
- specialised NICs with first class virtualisation support to avoid hypervisor overhead (mostly just a cloud thing)
One thing about high performance networking is that there are really two different desires: maximising throughput and minimising latency. Different choices will be made based on what kind of performance is most desired.
[1] if you imagine a typical high performance, pre-io_uring, portable web server, you have some socket which you listen on, many threads which call accept on it (with an option so that only one would be woken up at a time with new connections) and a load of fds for open connections with the server calling epoll to wait for inputs from them. And when a packet comes in and makes it through the kernel tcp stack, the kernel has to assign the packet to an appropriate fd and figure out if any thread needs to wake up. But instead with specialised networking configurations, you can just receive the packets in the order they arrive and send them off to their handlers yourself without the overhead of explaining how to do that to the kernel.
Intel even offers DDIO (Direct Data I/O), which allows an Intel NIC and CPU to share data without going through main memory.
DDIO applies to ANY I/O device, which makes it a horrible idea on servers with even a moderate amount of I/O. With DDIO, EVERY I/O device's DMA writes (eg, transfer from device to host) allocates space in cache. This means it works great on toy benchmarks with a single thread polling for I/O from a NIC. But scale it up to a 32c/64t server that is reading a few gigabytes/sec from disk and doing few tens of gigabits of network I/O, and you thrash your caches.
Its precursor, DCA, was limited to Intel (and a few other vendors) NICs, and was far better. With DCA, the NIC decides what, if any, DMA write should be pushed to cache and tags that DMA write with a PCIe TLP steering tag. The best choices are typically RX descriptors, and protocol headers. Other data is not tagged, and hence does not pollute the cache.
In terms of features like CAT being available on different SKUs.. Intel needs to remember that they're not just competing with themselves these days. Their competitors generally don't go fusing off features on low-end SKUs to drive sales of higher margin SKUs. Intel needs to differentiate chips solely on on cores count, clock speed, cache size, etc.
It's too bad that DDIO has not trickled down the product line. It doesn't exist on desktop-class parts, and many people can't use it even on top-of-the-line Xeon-SP because cloud operators disable it and RDT/CAT, which is helpful when using direct cache access.
Sometimes the entire frontend of a bank (Say: HSBC, which is a global bank) can fit into 3 racks of servers.
Not saying it's the right thing, just confirming your question.
I guess it will go into the JVM too.
I was just looking into building a custom router using an older PC and a dual-port 2.5GbE NIC. OpnSense sounds like a solid choice, but I'd prefer a Linux solution because the CPU is pretty beefy and wouldn't mind running a few Dockers on the side (yes, I know multi-purposing a router is a bad idea but meh).
Additionally, you could set the box up as a hypervisor, pass the NIC through to the router VM, and have the other VMs/containers use a bridge or some more advanced virtual networking.
That obviously has potential downsides, but I had a bit of fun doing it before downsizing my server
I didn't read the entire thread, but are there still implementation kinks that are being worked out or is the api ready to use now?
I went through and read it, and the submitter is incredibly confrontational and not at all open to feedback on the correctness of their benchmark. Also, when presented with contradictory evidence (their own benchmark where the results show io_uring is faster than epoll on other machines), they essentially dismiss it and says the other users ran their benchmark wrong, or that they can't reproduce their results on their machine, or that the other users used Boost and therefore are invalid. So not an entirely reliable criticism coming from them, in my opinion.
My understanding at the moment is that there is not one io-uring but every kernel version has a different implementation, and it changed quite a bit over the years. In the earlier version it just delegated all IO to an in kernel threadpool. Especially for network IO that’s not ideal, since the alternatives (epoll) didn’t require a threadpool neither in kernel nor userspace. And this probably showed up in benchmarks. The implementation changed, but I can’t really tell how now everything works for disk and network IO with out reviewing source code again.
A comprehensive changelog which explains implementation changes and bugs in different versions would be super helpful
Let's say they have a circular buffer for the TX queue. Why can't they have an atomic variable for the read end of the circular buffer; one could map this variable into shared memory, the kernel could then move this counter, once the nic has completed to send a frame with the data from the read end of that circular buffer. User space could still query this counter to check for 'buffer full' conditions.
Unfortunately the whole project was killed and all Fiber projects pulled from our office before we could finish.
Eliminating a kernel copy of the buffer being made will improve the performance significantly, but probably won't get you the same performance as a specialised card which would be what you'd usually go for if it's that important.
Generally speaking:
(A) prevent kernel from touching data and only operate on metadata (to reduce cache misses/memory bandwidth)
(B) in routing situations prevent application from copying data
(C) busy-polling - save cpu on doing everything in application as opposed to jumping to/from kernel context, for better latency
(D) allow manual buffer management
For example, splice or tcp_mmap() can be used to "receive" data from networking stack into application managed place. splice and MSG_ZEROCOPY can be used to "transmit" data into networking stack. The problem with MSG_ZEROCOPY is that, while buffers belong to application, the API doesn't tell you when it's done sending and it's possible to repurpose the buffers.
This io_uring article seem to indicate this fault has been fixed in io_uring zerocopy API.
These API's are pretty much (A), prevent kernel from touching data, and only look at metadata.
Due to API limitations it's hard or even impossible to do (B) in application. But with better API's we'll get there.
To get (C), you need raw hardware access basically.
But the holly grail is (D): the buffer-reuse problem. Ideally you'd want to submit buffer to networking card (or network stack), get data from NIC directly into there, then get a handle to this data in application, perform routing decision, submit that buffer for transmission in some other NIC, and eventually get a completion notification. At which point the buffer could be submitted for RX again. The userspace zerocopy API's are far away from this ideal scenario.
So no. Zerocopy kernel API's don't immediately solve all problems typical kernel bypass solution aims for. Having staid that, I don't think many users have the problems true kernel bypass solves.
This io_uring article seem to indicate this fault has been fixed in io_uring zerocopy API.
This is the hard part of a zero-copy API, and why its hard to make a zero-copy API that's a drop-in replacement for sockets and can magically enable naive applications to have zero-copy transmits. I attempted to do zero-copy sockets in FreeBSD ~20+ years ago (https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.34...). What's not mentioned in the paper is that the change was an overall net-negative, because applications would frequently trigger a COW fault when attempting to write to a buffer owned by the kernel.
i've been thinking about these sorts of things (kernel/user data barriers and high performance ipc) a lot lately as i've been wondering what a high performance, high security microkernel might look like.