Io uring
nick-black.com
nick-black.com
Our general take was also that it has a lot of potential, but is relatively low level that most mainstream programmers aren't going to pay attention to it. Hence, it'll be a while before it permeates through various ecosystems.
For those of you that like to listen on the way to work, we cover io_uring on our podcast, The Technium.
https://www.youtube.com/watch?v=Ebpnd7rPpdI
https://open.spotify.com/episode/3MG2FmpE3NP7AK7zqQFArE?si=s...
this may take a while as it's a completely different IO model
it took us 30 odd years to get from select/epoll to async/coroutines being popular
(which no-one uses either because the API doesn't compose well onto existing application structures)
But POSIX async I/O (the aio_* functions) in Linux is basically worthless performance-wise AFAIU, because Glibc implements it in userspace by spawning threads to do standard sync I/O. Now Linux also has non-POSIX async I/O (the io_* functions), but it’s very situational because it works only if you bypass the cache (O_DIRECT) and can still randomly block on metadata operations (so can Win32, to be fair). There’s select/poll/epoll with O_NONBLOCK of course, which is what people normally use, but those do not really work with files on disk (neither do their WinSock equivalents). Hell, signal-driven IO (O_ASYNC) exists, I’ve used it to make a single-threaded emulator (CPU-bound unlike a network server) interact with the terminal. But asynchronous I/O of normal, cached files is only possible on Linux through the use of io_uring, as far as I’ve been able to figure out.
That said, I’ve read people here saying[1] that overlapped I/O on Windows also works by scheduling operations on a thread pool, even referencing KB articles[2]. This does not mesh with everything I’ve read about I/O in the NT kernel, which is supposed to be natively async to the point where the I/O request datastructure (the IRP) has what’s essentially an emulated call stack inside of it, in order to allow the I/O subsystem to juggle continuations. What am I missing? Does the Win32 subsystem need to dumb things down that much even inside its own implementation?
(Windows 8 also introduced a ringbuffer-based, no-syscalls thing called Registered I/O that looks very much like io_uring.)
The _kernel_ thread pool. Eventually, most work has to be done in an actual thread, after all.
> [2] https://support.microsoft.com/kb/156932
It's a bit misleading. What they mean is that some operations can act as barriers for further operations. E.g. async calls to ReadFile won't run until the call to WriteFile finishes (if it's writing past the end of the file).
> O_NONBLOCK [...] has no effect for regular files and will (briefly) block when device activity is required, regardless of whether O_NONBLOCK is set. []O_NONBLOCK semantics might eventually be implemented[.]
I’m actually not sure if the reported readiness for them is of any use, but the documentation for select(2) [2] doesn’t give me a lot of hope:
> A file descriptor is ready for writing if a write operation will not block. However, even if a file descriptor indicates as writable, a large write may still block.
This for data operations; if you want open() itself to avoid spelunking through NFS or spinning up optical drives or whatnot, before io_uring you simply had no way to tell that to the kernel—you call open*() or perhaps creat(), which must give you a fd, thus must block until they can do so.
(As far as I’ve seen, tutorial documentation usually rounds this down to “you can’t do nonblocking I/O on disk files”.)
3 decades.
There's also a syscall (io_uring_enter) that will do a context switch and wake you up when completions are available (it's a complicated syscall, that has a lot of knobs and switches and levers - just be ready for a LOT of information if you go read the man page).
On a related note, I recently saw this presentation[1] where they show some benchmarks of the various modes.
One gotcha of sorts, though obvious when you think about it, is that the kernel-side polling mode requires a free CPU core for the polling thread. Meaning you'll get very poor performance if you're not leaving enough CPU for the kernel to do its polling.
Isn't that what libraries are for?
https://github.com/uNetworking/uWebSockets/issues/1603#issue...
A big todo I have for this is to modify it to accept multishot_accept but I haven't been able to get a recent enough kernel configured properly to do it. (Only >6.0 kernels support multishot accept).
(edit) you need to s/multishot_accept/recv_multishot/g above :)
Thank you so much Martin Jacob!
The things you work on are foundations to extremely useful systems.
(My JIT compiler is based on Martin Jacob's JIT code https://github.com/samsquire/compiler )
I would love to change my epoll server to use liburing and if I feel confident enough to I'll try integrate your code into it.
At least, that's my untested understanding.
“this is the wiki of nick black (aka dank), located at 33°46′44.4"N, 84°23'2.4"W (33.779, 85.384) in the heart of midtown atlanta. dankwiki's rollin' wit' you, though I make no guarantees of its correctness, relevance, nor timeliness. track changes using the recent changes page. I've revived DANKBLOG, this wiki and grad school having not satisfied ye olde furor scribendi.
hack the planet! don't mistake my kindness for weakness.
i primarily write to force my own understanding, and remember things (a few entries are actually semi-authoritative). i'm just a disreputable Mariner on your way to the Wedding. if you derive use from this wiki, consider yourself lucky, and please get confirmation before relying on my writeups to perform surgery, design planes, determine whether a graph G is an Aanderaa–Rosenberg scorpion, or feed your pet rhinoceros. do not proceed if allergic to linux, postmodern literature, nuclear physics, or cartoonish supervillainy. “
I’ve had his grad-school project libtorque[2] (HotPar ’10), an event-handling and scheduling library, on my to-read list for years, but I can’t seem to figure out how it accomplishes the interesting things it does.
[1] https://nick-black.com/dankwiki/index.php/Notcurses, https://github.com/dankamongmen/notcurses/
1. enroll.
2. get tired of dorm life.
3. move to home park because it's likely all you can afford.
4. get broken into 5-10 times.
5. move away
Not one break in!
Figured no one thought we were worth it. Although we were two houses across the street from Kool Korner, so maybe that was the "nicer" part of Home Park?
I should have been more precise re: break-ins, meaning more car break-ins. No idea what the home rate was, but every other week there were 1-2 cars missing windows around Lynch.
The west & north sides of GT are unrecognizable to me now though -- condos and yuppie shops.
I'd be curious to see
1. A Hello World, absolute minimum example
and
2. Doing something it's designed to do well, cut down to as small an example as possible.
Sheez, downvoters, I'm curious about it, and want to see an example. You don't learn to drive a car by reading the engine specs, either.
Right now, what it can do really well is non-blocking file IO. My (limited) understanding is that as of now, the benefits of io_uring over epoll for network IO is a bit more ambiguous. That said, io_uring is adding new features (already available in linux 6 kernel) that are really promising. See https://github.com/axboe/liburing/wiki/io_uring-and-networki....
https://www.scylladb.com/2020/05/05/how-io_uring-and-ebpf-wi...
Look at AMD gpu command buffers, and xHCI (and NVMe).
But among those "ring buffers", which one is the best? (from a consumer/producer concurrent access, on a modern CPU ofc).
If the commands are not convoluted, the programming is soooo much simpler and cleaner.
It seems USB xHCI is the "most concurrent friendly" and hardwarely friendly (IOMMU and non-IOMMU) as it supports "sparse ring buffers" (non continous in bus address space). AMD gpu ring buffers are atomic read/write pointers (command ring buffers and "interrupt"/completion ring buffers) with doorbells to notify pointer updates.
I should have a look at linux IO uring to see how far they did go and which model they did choose.
I wonder how modern hardware could use anything else than command ring buffers.
Including a tl;dr, pros, cons, tools & players, and a forecast.
The creator of io_uring Jens was super nice and also reviewed it (I'm still a bit in shock).
Pretty cool to see how far IO APIs have come with this.
Glommio is a thread-per-core async executor/framework that supports io_uring however.
I read a while back that it only did fs, that was clearly incorrect/outdated/etc.