Libuv – Linux: io_uring support
github.com
github.com
Since you can't tell if a disk io function would have to actually hit the disk or use the cache on older linuxes, libuv just always dispatches said operations to a worker thread pool which then calls the normal function. When it returns, it sends a message back to the other thread to indicate its completion.
If data is already present in the cache, this adds significant overhead vs just calling `read` directly - instead calling `read`, it will post a request to the worker thread queue (which may require synchronization, im not sure exactly how the impl does it tbh), wake the worker thread which calls `read`, which does not block significantly since the data is indeed already present, then it posts an event back to the main thread (which, sure, it is listening for this event on `epoll`, along with whatever else it is listening to, but that's an insignificant detail at this point) which calls the on-complete handler. All that work instead of just transferring the data from the cache to the result. Significant slowdown.
If the data is not already present in the cache, it does exactly those same steps, except this time `read` actually does block for some time, so the main thread can continue with other work in the mean time. (Of course, if the main thread has no other work to do, you've gained nothing from this!)
Hence why I said in my other comment that you ought not draw any general conclusion from this. This speedup has nothing to do with epoll vs io_uring (except maybe that linux file i/o is not really epoll compatible) and everything to do with libuv's block device implementation specifically. It is totally inapplicable to network loads entirely.
edit: in the first version I said "disk file" but it is technically block devices, of which disk files are just the most common example, but of course /dev/zero is a block device that is not a disk file.... and worth noting that /dev/zero is never going to actually load off a disk meaning it is "in the cache" all the time.
A couple years ago I have added io_uring to my toy webserver but it was slower than epoll. I would love to see someone using io_uring for a real networking application and report more meaningful statistics than microbenchmarks.
99.9% of the time the data will not be evicted prior to the read, so it's still going to be a win overall.
(If you haven't read about Windows' overlapped i/o, do so, it is actually quite nice to use.)
io_uring brings something similar to Linux.
The nonblocking read just failed. The task initiates IO then because the IO is async it lets other tasks run while waiting for data. Later, when the IO finished, the task can read the memory.
Is this what you mean?
I think there must be a data race somewhere or io_uring would be superfluous. Everything consisting of at least two interruptible steps at the lowest level in the CPU but where people expect no change in state is subject to some data race. It's really difficult to get this 100% correct. Something works fine for 99.99% of the time. Such a situation can be very nasty to get it right.
Perhaps the step "Later, when the IO finished, the task can read the memory" cannot be done atomically.
If you're using io_uring, the data is read into the user buffer by the kernel during the async call, so there's no race. There's no point having a separate fast path in that case, because the io_uring call can just return the data immediately if it's available.
Question I’ve not figured out yet: how can one trace io_uring operations? The api seems kinda incompatible with ptrace (which is what strace uses) but maybe there is an appropriate place to attach an ebpf? Or maybe users of io_uring will have to add their own tracing?
which is nice because you can measure the actual library overhead without the noise from real kernel work.
I would think that writing the kernel part would be the hardest, but it's usually the event loop implementations that don't use what the Windows/MacOS/Linux kenels offer.
To give an example, libnbd does a lot of send and recv system calls interleaved, for sending requests and receiving replies, with a complicated state machine to keep track of when we would block. It's still not obvious how to convert this to an io_uring style where you have to add the send/recv requests to the submission queue and then separately pick up the responses from the completion queue. It's a particular problem dealing with possible errors and keeping the API the same (if you just call exit when you hit any error, it's a lot easier).
If you're writing all new code then io_uring is a great choice.
I know it's a different library, and asyncio is not as flexible as io_uring, but at least batching send / recv calls from different sockets would be good to be able to do.
Everyone is adopting it pretty quickly, actually.
At least the .NET entry is however no longer correct
From what I understand, some async operations will be faster in node on newer versions of the linux kernel, when node uses a version of libuv that contains this PR.
For example, implementing tcp/http/udp proxy servers. Static file servers too can benefit greatly.
> is this PR just for files or also for network operations?
> Just file operations for now.
Not sure, why they ignored socket, it would be have been great addition to NodeJS/Python stack. Let's hope they make it happen in near future release schedule.
In other words: if your application is already writing/reading at the max speed the hardware its on allows, you will see no improvement from this. That application is truly I/O bound.
An application that feels I/O bound will be doing lots and lots of little reads and writes and achieving low hardware utilization. That application is syscall bound rather than I/O bound, and io_uring will help tremendously.
Otherwise, i.e. in the real world, it's very unlikely to be that much. High throughout workloads should see improvement, but each application is different.
Github post does it normally: 'Add io_uring support for several asynchronous file operations:'
Explaining with examples:
* [io_uring support] [for libuv]: "libuv has io_uring support"
* [io_uring] [support for libuv]: "io_uring has support for libuv"
* [Windows subsystem] [for Linux]: "Linux has a Windows subsystem"
* [Windows] [subsystem for Linux]: "Windows has a subsystem for Linux"
[x] [y for z] / "x has y for z" is much more common.
They couldn't even be bothered to put an apostrophe on the end of Windows.
1. I think they are not the same senses. In ‘windows subsystem for Linux’, the word means ‘having as a function’ like in the phrase ‘spanner for 1/4 inch hex bolts’.
2. I think the sense in the title is more like ‘to the benefit of’ (or perhaps ‘affecting’ or ‘having the reason’) like ‘lunch for employees’ or ‘supports for the lintel’
3. I checked a couple of dictionaries and they had the sense but definitions for words like these can be pretty hard to read even for a native speaker.
For what it's worth, I'm a native English speaker and I agree they both sound the wrong way round to me! But I can convince myself that they do also make sense the way round they were intended.
(in case the title changes later, it's currently "io_uring support for libuv")
You make a really good point I never heard before: X for Y is an ambiguous construct with two names referring to the same class of object, without an obvious relationship. Ex. Barack for Don is easy, but only if you know American politics, a reference to “a predecessor”. But “Jon for Mary” is inscrutable.
In a more normal construction ("food voucher support for kids") it is obvious from context that the kids are being given the support, because the converse would be nonsensical.
So, by that logic, clearly then Windows Subsystem for Linux is a version of Windows specifically suited for running under Linux.
"xyz for Windows" is allowed, while "Windows xyz" is reserved by the company to indicate that it is actually an official part of Windows.
If you haven't yet, please go check it out, write a program with it and be amazed.
So glad to be a contributor.
To me libuv seems highly rated and it deserves to be highly rated. I'm not sure how to quantify it exactly to say whether it's fairly rated, underrated, or overrated, but it seems like it's in the ballpark as far as ratings go.
Also, a polling loop + how to properly use it should be something as fundamental as pointers, linked lists, ..., all that stuff.
That indeed makes it underrated. Got it.
Without libuv that story would not have been such a success.
Thank you!
Is that a very good thing or a very bad one?
I could be mistaken though, it seems odd not to post any sort of numbers/benchmarks.
The commenter mentions "greater than 8x" on a very artificial benchmark of reading /dev/zero. The original PR author than copies that to the PR description as "8x has been observed" (lossily dropping the greater, but not adding any additional datapoints).
The reality will be that disk IO dwarfs the overhead of syscalls and threadpools in all real cases.
That all that's being saved here: the overhead of having to manage a threadpool, and make a few more syscalls.
Compared to reading a few MB off disk, that overhead will not be noticed, and definitely won't be 8x.
Does this mean libuv already supported io_uring for non-file operations? Or it still doesn't?
Async file operations are useful in some applications, but not the main things people normally think of when they hear async IO.
Just one questions: what about older versions of linux that don't have io_uring, does it fall back gracefully to older system calls or are these older versions of linux no longer supported?
This isn't to say that io_uring is bad, just don't draw too much a conclusion from any benchmark of their old impl beyond the context of their old impl specifically.
I've been studying how to create an asynchronous runtime that works across threads. My goal: neither CPU and IO bound work slow down event loops.
How do you write code that elegantly defines a state machine across threads/parallelism/async IO? How do you efficiently define choreographies between microservices, threads, servers and flows?
I've only written two Rust programs but in Rust you presumably you can use Rayon (CPU scheduling) and Tokio (IO scheduling)
I wrote about using the LMAX Disruptor ringbuffer pattern between threads.
https://github.com/samsquire/ideas4#51-rewrite-synchronous-c...
I am designing a state machine formulation syntax that is thread safe and parallelises effectively. It looks like EBNF syntax or a bash pipeline. Parallel steps go in curly brackets. There is an implied interthread ringbuffer between pipes. It is inspired by prolog, whereby there can be multiple conditions or "facts" before a stateline "fires" and transitions. Transitions always go from left to right but within a stateline (what is between a pipe symbol) can fire in any order. A bit like a countdown latch.
states = state1 | {state1a state1b state1c} {state2a state2b state2c} | state3
You can can think of each fact as an "await" but all at the same time. initial_state.await = { state1a.await state1b.await state1c.await }.await { state2a.await state2b.await state2c.await } | state3.await
In io_uring and LMAX Disruptor, you split all IO into two halves: submit and handle. Here is a liburing state machine that can send and receive in parallel. accept | { submit_recv! | recv | submit_send } { submit_send! | send | submit_recv }
I want there to be ring buffers between groups of states. So we have full duplex sending and receiving.Here is a state machine for async/await between threads:
next_free_thread = 2
task(A) thread(1) assignment(A, 1) = running_on(A, 1) |
paused(A, 1)
running_on(A, 1)
thread(1)
assignment(A, 1)
thread_free(next_free_thread) = fork(A, B)
| send_task_to_thread(B, next_free_thread)
| running_on(B, 2)
paused(B, 1)
running_on(A, 1)
| { yield(B, returnvalue) | paused(B, 2) }
{ await(A, B, returnvalue) | paused(A, 1) }
| send_returnvalue(B, A, returnvalue)