Also, I think you misunderstood the blog as it's describing application read-ahead which is what you have to do when using O_DIRECT.
preadv2 + RWF_NOWAIT
There really aren’t optimal user space APIs. What you want ideally is a way to register a callback on a memory region so that the kernel is able to treat it like page cache and then on a fault you regenerate it using said callback. Thats even more efficient because usually you have to do some processing of the on disk representation (eg if it’s compressed).
I implement read-ahead in the application by (optionally) preadv:ing a single read into multiple destination buffers in the pool, leaving them unpinned, since as long as you aren't up against the bandwidth limit of the drive, a larger read is generally as fast as multiple smaller one on modern hardware.
I've tried doing this with io_uring as well, but found just eating the preadv syscall cost was faster.
Not always if by modern you mean NVMe drives. One synchronous preadv() for 256 KiB gives the kernel/device one big request but 16 independent asynchronous 16 KiB reads can be serviced concurrently. So the latter gives the NVMe controller 16 operations it can schedule in parallel. So depending on the workload and hardware, offsets, filesystem and request sizes that can give you lower aggregate latency or higher throughput.
Modern SSDs tolerate moderate queue depths very well, but piling on the I/O queue also incurs tail latency jitter unless you're able to ensure the queue depth stays in the moderate range and never goes higher. All else being equal, fewer larger requests is better for I/O latency (though read amplification for the sake of reading more data obviously doesn't help anyone). Though in this scenario, we're mostly comparing the syscall overhead of a single preadv against io_uring bookkeeping for multiple preads, regardless of how you submit the reads they end up being the same operation.
In a poorly written micro benchmark test it’s possible the kernel will fail to coalesce all of them because your submissions aren’t visible all at once (ie it starts submitting requests and doesn’t have an opportunity to merge). Whether that actually is possible to happen requires digging a bit more into the Linux kernel source.
You can look at 'rrqm/s' and '%rrqm' in `iostat -x` to get statistics for how many requests are merged, and `rareq-sz` for the average request size.
In fact, one consistently sees higher bulk IO numbers when using physical media that has been formatted with a large sector size compared to the old 512 byte fixed emulated size. You’d routinely see lower latency and higher IOPs with 4kn (HDDs or SSDs) than you would with 512e disks, even with SCSI or AHCI controllers that featured similar pipelining support to today’s NVME controllers (or even if you place a spinning rust HDD behind NVMe today!).
That’s the point I was ultimately making.
Luckily the read heads only have one degree of freedom, so the "elevator algorithm" is sufficient: https://en.wikipedia.org/wiki/Elevator_algorithm
Replacing single sycalls with their equivalents in io_uring is generally slower than just making the syscall directly. io_uring still uses syscalls after all.
io_uring generally only wins if you can amortize its overhead across multiple simultaneous operations. Implementing readahead would be such a case, except you can accomplish the same amortization with a single preadv instead, which again turns into a single syscall for multiple reads.
The fair comparison isn’t 1 syscall on a single thread processing 1 task against io_uring. That would be insane because you clearly don’t have any performance requirements in such a workload already.
The closest realistic equivalent would be using Tokio’s spawn_blocking to do that 1 syscall vs doing that syscall in io_uring. It’s probably still more efficient if your benchmark literally is the cost of 1 syscall at a time but not by as much and io_uring in poll mode doesn’t even enter the kernel so it can actually outperform the syscall offloaded to a background thread (even though yes under the hood it’s the same kernel code).
Requests are also not I/O bound - you have a mix of CPU work to figure out where to issue the I/O and to manage the cache buffers on completion and read whatever information you need from the buffer pulled in.
But sure, if all your DBMS is going to be doing is one point lookup at a time then there’s no benefit. But even non trivial SQL queries will surprise you because you can get complicated IO famous (eg checking the index which then tells you multiple blocks to look through which you can do concurrently).
TLDR: unless your DBMS is ass slow, pread isn’t sufficient to get good throughput. And at that point it’s irrelevant if a single preadv is faster than io_uring.
[1] Mandatory mmap=poop-emoji link: https://db.cs.cmu.edu/mmap-cidr2022/
Its a new web server i am building and its the fastest way i could find out.
Just switching from epoll to liburing made the server ~45% faster too, its ridiculous. It can serve 10 gigabyte per second with a single thread, or around 10 million responses per second with h2 and 32 multiplexed requests.
I had to write a new http load generator for that since i couldn't find one which could generate enough load to saturate my server or be fast enough to withstand it.
He's very clearly hauling data straight from the page cache.
i open the file, mmap it, close it and tell io_uring_prep_send which bytes from the mapping to send, this also saves me from a possible SIGBUS cause the access to the mmap happen inside the kernel, and when a SIGBUS would happen in userspace the kernel just reports a shorter send in cqe->res
In my opinion, the POSIX advices specified for madvise are useless or even dangerous.
On Linux, for precise control of memory-mapped files one should use only these 4 Linux-specific advices: MADV_COLD, MADV_PAGEOUT, MADV_POPULATE_READ & MADV_POPULATE_WRITE.
These should be used within io_uring, so that they will be executed asynchronously.
These have a well-documented meaning and using them carefully should be sufficient to reach optimum performance with mmap.
mmap and io_uring_prep_send were faster than everything else, no matter the size as long as you keep the map around for the lifetime of the process.
for one off sends when a file is smaller than 256kb then io_uring_prep_read + prep_send are faster than everything else.