The Rapid Growth of Io_uring
lwn.net
lwn.net
(Incidentally, this is a good illustration of the flexibility of UNIX’s model of describing everything with a file descriptor. The same interface meant for asynchronous file I/O was easily extended to network I/O.)
That sounds unbelievably annoying.
No, the limit has been found and it is not in the kernel. You push the I/O loop up for userspace I/O, more specific, not more general. They call it SPDK [1] (or DPDK for networking) and as far I can tell, the principle is essentially having a dummy driver in the kernel that maps the entire PCIe peripheral memory space into your chosen process, and everything flows from there.
At the I/O limit, asynchronous isn't feasible because interrupts introduce latency and waste cycles not doing work. All userspace I/O frameworks work only through polling.
io_uring supports polled i/o: https://lore.kernel.org/linux-block/20190116175003.17880-8-a...
Userland packet processing (in network context) is much more flexible and less brittle than forcing certain functionality to exist solely in the Kernel layer. However things do exist that allow you to (mostly) transparently re-jigger a standard app's TCP/IP calls. One such example is using LD_PRELOAD to "hijack" the sys-calls for certain things and snake it to your (super high performance) userspace app!
There's a lot of exciting stuff happening in the open source networking world (DPDK, VPP/FDio, Network Service Mesh, etc). I really recommend digging into it!
(I say "relatively sane" because this isn't really asynchronous, it's just make-believe with a kernel-managed thread pool, because I/O being a fully synchronous affair is ingrained far too deep into both Linux and BSD I/O stacks.)
What's your criteria for "really asynchronous"?
> it's just make-believe with a kernel-managed thread pool
To the extent that io_uring uses anything resembling a thread pool, it seems to me that it is used completely differently from how a userspace AIO thread pool operates. When a userspace AIO implementation submits IO to the kernel, it does so with a blocking syscall and that thread stalls until that IO is complete. That means the number of outstanding IOs is limited by the number of threads in the pool. I don't see any such limitation in using io_uring to deliver IO requests to the block layer.
One example might be that related operations, like stat(), opendir(), readdir(), getpeername(), and so on...remain synchronous. And that async functionality is mostly a bolt-on to very established things, file descriptors, berkeley sockets etc.
Also, every improvement is a pretty big patchset with code to ensure the traditional synchronous operations don't get unintended side effects.
Just generally the idea that a "clean start", non-POSIX bound OS might design things differently, I imagine. Google's Fuchsia seems to hit some middle ground, where async is more foundational, for example.
I don't think that observation detracts from the improvements.
Microsoft's (research, discontinued) Midori OS was heavily async:
You can already have asynchronous IO from userspace to the block layer via linux-aio and O_DIRECT. But the VFS remains synchronous, so both uring and userspace can effectively only use a thread pool to work around that. Or magic.
The cost of syscalls has gone up massively due to all the mitigations against recent side channel attacks. Batching and async syscalls would be an even bigger win than ever.
Wouldn't it be great if glib or something could adopt new mechanisms and everyone got faster?
The dynamics and mechanical sympathy of different approaches are changing as hardware goes massively parallel faster than software can.
You'd still need a syscall to sleep when the ring is empty, and a syscall to wake up the mechanism when the first request is put on an empty ring. But yeah, other than that (and perhaps an "yield" syscall), you could do everything through the ring.
If the scale of the problem changes, so does the solution. We know this.
https://www.freebsd.org/doc/handbook/linuxemu.html
I wonder though... wouldn't a safe first way of implementing these new syscalls be to make them actually synchronous?
That way you'd be able to run these Linux binaries but without any of the performance benefits.
No, because it visibly changes the semantics. Consider for instance IORING_OP_ACCEPT; if you make it synchronous, and nothing connects to your program, it would wait forever, instead of returning immediately and allowing the program to continue. The file-related opcodes are safer (when used with actual files, instead of network sockets), but still would behave differently for instance with a hanging NFS mount.
You mean thread-pools for file I/O?
From the abstract:
"iouring enhances the existing Linux AIO API, and provides QEMU a flexible interface, allowing you to use the desired set of features: submission polling, completion polling, fd and memory buffer registration. By explaining these features we will come to examples of how and when you need to use them to get the most out of iouring. Expect many benchmarks with different QEMU I/O engines and userspace storage solutions (SPDK).
You will get a brief overview of the new kernel feature, how we used it in QEMU, combined its capabilities to speed up storage in VMs and what performance we achieved. Should io_uring be the new default AIO engine in QEMU? Come and find out!"
[1] io_uring in QEMU: high-performance disk I/O for Linux — https://fosdem.org/2020/schedule/event/vai_io_uring_in_qemu/
eventually libraries like Libuv turned up and made my life a lot easier.
Are there any good stats on how io_uring compares to all those older async io stuff?
Don't know how it compares to epoll performance-wise.
In fact io_uring might have direct support for polling now.
There are several io_uring opcodes intended specifically for the network use case, including: IORING_OP_ACCEPT, IORING_OP_CONNECT, IORING_OP_SENDMSG and IORING_OP_RECVMSG.
Did you mean something else by 'target'?
io_uring: https://github.com/frevib/io_uring-echo-server
epoll: https://github.com/frevib/epoll-echo-server
kqueue: https://github.com/frevib/kqueue-echo-server
io_uring performs much better than epoll: https://twitter.com/hielkedv/status/1218891982636027905?s=21
Note that it might have been optimized more in the meantime.
I've seen database programmers rave about the speed benefits of asynchronous I/O, because they have to store a lot of data on disk. But the majority of programs I write only have to deal with reading files. I'd love to try using io_uring, but only when it's appropriate.
If your program needs to read in an entire JSON blob (for example) before it can do anything, or if it only does light processing on each individual part (like adding up a column of a CSV file) then async I/O probably isn't going to help.
You can also build web servers that read from disk as part of the request handling, without having to do blocking file reads. Before io_uring, this scenario would get absolutely horrible performance as the whole thread would lock down while waiting for the file read. Now you can do other things concurrently.
SELECT client_id, invoice_id, invoice_status, invoice_date
FROM invoices
WHERE client_id IN (123, 567) AND invoice_status != 'PAID';
Imagine that there is an index for client_id.The DB engine can scan the index, finding all the pages in the table that contain rows with one of those two client_ids in.
It requests these pages as soon as it finds them in the index.
It scans the pages in the order it finds them completed, though, which can be a completely different order than it requested them in.