1) create one io_uring per thread (much like you'd create one epoll grp)
2) use the provided buffer mechanism to post a pool of large buffers to an io_uring. Bonus points for the newer ring based version.
3) keep a large (128k) async multishot recv posted to every socket in your set always
4) as recvs complete, append the next "segment" of the stream to a per-socket list.
5) parse the protocol stream. As you make it through each segment, return it to the pool*
6) aggressively batch data to be sent. You can only have one outstanding at a time per socket, so make it a big vectored write. Writes are only outstanding until they're confirmed queued in your local kernel, so it is a fairly short time until you can submit more, but it's worth batching into a single larger operation.
* If you need part of the stream to live for an extended period of time, as we do for the payloads in NVMe-oF, build scatter gather lists that point into the segments of the stream and then maintain a reference counts to the segments. Return the segments to the pool when it drops to zero.
Everyone knows the best way to use epoll at this point. Few of us have really figured out io_uring. But that doesn't mean it is slower.
seastar.io is a high level framework that I believe has "figured out" io_uring, with additional caveats the framework imposes (which is honestly freeing).
Additionally the rust equivalent: https://github.com/DataDog/glommio
As an aside, it's great to see high praise from an spdk maintainer. One of the big reasons for doing io_uring in the first place was that it was impossible to compete in terms of performance with total bypass unless you changed the syscall approach.
io_uring was only a little slower, and there are some advantages to io_uring with regard to adaptive performance (because Linux doesn't expose some information to userspace that's useful for this, so userspace has to estimate with lag - see Go's scheduler), but I was hoping it would be significantly faster. Then again it was good to have an alternative to validate the thread pool design.
Of course, batching is also great.
Windows 11 copied io_uring almost 1:1 - https://windows-internals.com/ioring-vs-io_uring-a-compariso...
The only way to avoid big latencies on uring_enter is to use the submission queue polling mechanism using a background kernel thread, which also has its own set ofs pro's and con's.
I don't have much of a feel for this because I am on the "never calling io_uring_enter" plan but I expect I would have found it alarming if it took 1ms while I was using it
There are certainly applications that don't benefit from io_uring, but I suspect these are not the norm.
Eg a UDP syscall will do a whole lot of route lookups, iptable rule evaluations, potential eBPF program evaluations, copying data into other packets, splitting packets, etc. I measured this to be fare more than > 10x of the syscall overhead. But your mileage might vary depending on which calls you use.
As for the applications: these lessons where collected in a CDN data plane. There’s hardly any applications out there which are more async IO intense.