Io_uring Without Readahead
frn.sh
frn.sh
This isn’t just with regards to benchmarking, it’s an essential consideration *any* time you are taking ownership of the cache away from the kernel, which is the only piece in the stack that has viability into the global state and can be trusted to give back memory under pressure to ensure everything plays nice together. It’s not limited to just databases or even just memory, for example FreeBSD has had greater than its fair share of issues that trace back to the ZFS having a separate cache from the kernel (despite the much tighter integration between the two and presence of various mechanisms to address pathological cases). You can also refer to any comparison or benchmark between the use of spinlocks vs mutexes: spin locks consistently perform better in/on (micro)benchmarks but are almost always actually the worse choice in the grand scheme of things because the benchmarks falsely assume complete and uncontended ownership over system resources.
This isn’t even io_uring specific and I’m hardly the first to bring this up in the context of O_DIRECT.
Also, I think you misunderstood the blog as it's describing application read-ahead which is what you have to do when using O_DIRECT.
I implement read-ahead in the application by (optionally) preadv:ing a single read into multiple destination buffers in the pool, leaving them unpinned, since as long as you aren't up against the bandwidth limit of the drive, a larger read is generally as fast as multiple smaller one on modern hardware.
I've tried doing this with io_uring as well, but found just eating the preadv syscall cost was faster.
Not always if by modern you mean NVMe drives. One synchronous preadv() for 256 KiB gives the kernel/device one big request but 16 independent asynchronous 16 KiB reads can be serviced concurrently. So the latter gives the NVMe controller 16 operations it can schedule in parallel. So depending on the workload and hardware, offsets, filesystem and request sizes that can give you lower aggregate latency or higher throughput.
Modern SSDs tolerate moderate queue depths very well, but piling on the I/O queue also incurs tail latency jitter unless you're able to ensure the queue depth stays in the moderate range and never goes higher. All else being equal, fewer larger requests is better for I/O latency (though read amplification for the sake of reading more data obviously doesn't help anyone). Though in this scenario, we're mostly comparing the syscall overhead of a single preadv against io_uring bookkeeping for multiple preads, regardless of how you submit the reads they end up being the same operation.
In a poorly written micro benchmark test it’s possible the kernel will fail to coalesce all of them because your submissions aren’t visible all at once (ie it starts submitting requests and doesn’t have an opportunity to merge). Whether that actually is possible to happen requires digging a bit more into the Linux kernel source.
You can look at 'rrqm/s' and '%rrqm' in `iostat -x` to get statistics for how many requests are merged, and `rareq-sz` for the average request size.
In fact, one consistently sees higher bulk IO numbers when using physical media that has been formatted with a large sector size compared to the old 512 byte fixed emulated size. You’d routinely see lower latency and higher IOPs with 4kn (HDDs or SSDs) than you would with 512e disks, even with SCSI or AHCI controllers that featured similar pipelining support to today’s NVME controllers (or even if you place a spinning rust HDD behind NVMe today!).
Luckily the read heads only have one degree of freedom, so the "elevator algorithm" is sufficient: https://en.wikipedia.org/wiki/Elevator_algorithm
That’s the point I was ultimately making.
Replacing single sycalls with their equivalents in io_uring is generally slower than just making the syscall directly. io_uring still uses syscalls after all.
io_uring generally only wins if you can amortize its overhead across multiple simultaneous operations. Implementing readahead would be such a case, except you can accomplish the same amortization with a single preadv instead, which again turns into a single syscall for multiple reads.
The fair comparison isn’t 1 syscall on a single thread processing 1 task against io_uring. That would be insane because you clearly don’t have any performance requirements in such a workload already.
The closest realistic equivalent would be using Tokio’s spawn_blocking to do that 1 syscall vs doing that syscall in io_uring. It’s probably still more efficient if your benchmark literally is the cost of 1 syscall at a time but not by as much and io_uring in poll mode doesn’t even enter the kernel so it can actually outperform the syscall offloaded to a background thread (even though yes under the hood it’s the same kernel code).
Requests are also not I/O bound - you have a mix of CPU work to figure out where to issue the I/O and to manage the cache buffers on completion and read whatever information you need from the buffer pulled in.
But sure, if all your DBMS is going to be doing is one point lookup at a time then there’s no benefit. But even non trivial SQL queries will surprise you because you can get complicated IO famous (eg checking the index which then tells you multiple blocks to look through which you can do concurrently).
TLDR: unless your DBMS is ass slow, pread isn’t sufficient to get good throughput. And at that point it’s irrelevant if a single preadv is faster than io_uring.
preadv2 + RWF_NOWAIT
There really aren’t optimal user space APIs. What you want ideally is a way to register a callback on a memory region so that the kernel is able to treat it like page cache and then on a fault you regenerate it using said callback. Thats even more efficient because usually you have to do some processing of the on disk representation (eg if it’s compressed).
[1] Mandatory mmap=poop-emoji link: https://db.cs.cmu.edu/mmap-cidr2022/
In my opinion, the POSIX advices specified for madvise are useless or even dangerous.
On Linux, for precise control of memory-mapped files one should use only these 4 Linux-specific advices: MADV_COLD, MADV_PAGEOUT, MADV_POPULATE_READ & MADV_POPULATE_WRITE.
These should be used within io_uring, so that they will be executed asynchronously.
These have a well-documented meaning and using them carefully should be sufficient to reach optimum performance with mmap.
Its a new web server i am building and its the fastest way i could find out.
Just switching from epoll to liburing made the server ~45% faster too, its ridiculous. It can serve 10 gigabyte per second with a single thread, or around 10 million responses per second with h2 and 32 multiplexed requests.
I had to write a new http load generator for that since i couldn't find one which could generate enough load to saturate my server or be fast enough to withstand it.
He's very clearly hauling data straight from the page cache.
i open the file, mmap it, close it and tell io_uring_prep_send which bytes from the mapping to send, this also saves me from a possible SIGBUS cause the access to the mmap happen inside the kernel, and when a SIGBUS would happen in userspace the kernel just reports a shorter send in cqe->res
mmap and io_uring_prep_send were faster than everything else, no matter the size as long as you keep the map around for the lifetime of the process.
for one off sends when a file is smaller than 256kb then io_uring_prep_read + prep_send are faster than everything else.
O_DIRECT is to be used with files that will be accessed in a random order (random meaning that the order is not predictable by the kernel) and with buffer sizes per access great enough that you want to avoid their copying between kernel and application (e.g. at least a few kByte per access).
Whenever O_DIRECT is used, the programmer takes responsibility to implement an adequate form of read-ahead, based on the access pattern that is predicted for the application, and which cannot be guessed by the kernel.
If the programmer did not implement read-ahead, like it was the case before doing the update described in TFA, that was an incompletely written program. One should not enable O_DIRECT without a complete implementation for its requirements, as that can lead only to lower performance than standard I/O.