Axboe Achieves 8M IOPS Per-Core with Newest Linux Optimization Patches
phoronix.com
phoronix.com
Nevertheless, io_uring is certainly the better design and I'm happy to see progress still being made.
AFAIK as I know, these changes shouldn't negatively impact tasks that are not making any use of these features. So, even if improvements are not directly proportional to what is being reported, they're still real.
Historically the benchmarks have been crafted in a way to make them seem faster than their real-world and well-formed opponents (e.g. the io_uring vs epoll benchmark). One issue that was pointed out was that the io_uring benchmark eschewed proper error handling in order to avoid some branches, whereas the epoll benchmark properly error checked. This reduced the benchmark times considerably, though no correction was ever published after the fact.
This is why in the Github issue I mentioned in another comment there are some people understandably a bit annoyed. I don't choose to jump to conclusions about Axboe's intent, and my original comment wasn't meant to do so.
The person you’re going after here probably felt compelled to counter the needless personal nature of your remarks. It’s difficult to experimentally verify relativity but we don’t criticize Einstein as a result.
Could you provide any links to these discussions?
Original claims were in the 90% and above performance increase over epoll. Then issues were found, and the figure was adjusted to 60% over epoll. Then more issues were found, and now real-world performance tests are showing minimal speedups if any.
Unfortunately the sibling commentors don't see "computer science" as a science but instead as a "feel good hobby", it seems. My point wasn't to hurt feelings, it was to provide a word of caution with these sorts of groundbreaking claims with respect specifically to the io_uring efforts, as they have been disingenuous quite a few times historically.
I don't doubt Jens does fantastic work. I don't doubt that he's seen these speedups in very specific cases. But people are celebrating this as a win where they'd be skeptical of e.g. "breakthrough" treatments of cancer (footnote: in mice). It's the same thing.
That's literally the entire point I'm making.
Someone claimed that you don't get to see the benefit of io_uring in a hypervisor[1], but they did not provide benchmark results.
[0]: https://github.com/axboe/liburing/issues/189#issuecomment-94...
[1]: https://github.com/axboe/liburing/issues/189#issuecomment-73...
[1] https://github.com/netty/netty/issues/10622#issuecomment-701...
Assuming you have an hardware DMA ring buffer for I/O that is directly mapped in userspace, the only thing that is really needed is to write the operation type, size, disk position and memory position to the ring buffer, update the buffer position and check for flush, doable in around 8 CISC instructions (plus the slowpath), so around 100x inefficient.
Without the direct mapped ring buffer and with a filesystem, you need a kernel to translate from uring to the hardware ring buffer, and here it still seems around 10x inefficient as around 100 instructions should be enough to do the translation (assuming pages already mapped in the IOMMU, that you have the file block map in cache, and that the whole system is architected to maximize the efficiency of this operation).
This type of thing is not measured in instructions anymore, it's measured in off-board operations and latencies. Cache misses, latencies to poke MMIO and get back interrupts if necessary, DMA transfer to complete, device access time. In this case it seems the hardware is theoretically capable of about 12M so the core mostly be just waiting for that.
For some workloads it's not that hard to have that deep queues. What's harder is to know when to use them and when not. There's really not enough information available to make any of this self-tuning.
EDIT: 512, looking at the screenshot..
Userspace applications that already have some support for asynchronous disk IO (either through the old libaio APIs or as a cleanly-abstracted thread pool) should be able to switch to using io_uring as their backend without too much trouble, and reap the benefits of async that actually works reliably (if switching from libaio) and with vastly lower overhead (if switching from a thread pool). Databases like PostgreSQL were some of the few applications that attempted to deal with the limitations of libaio, but I'm not sure how close they are to having a production-quality io_uring backend.
"His patches pushing the greater performance have been changes to the block code, NVMe, multi-queue blk-mq, and IO_uring." https://git.kernel.dk/cgit/linux-block/log/?h=perf-wip
So it looks like he is playing with a good portion of the IO block stack with a very recent concentration on io_uring. So maybe some of it?...
edit- It's someone's name, I thought it was a company or product
[1] https://git.savannah.gnu.org/cgit/coreutils.git/commit/?id=2...
There are also optimizations on other layers, e.g. transparent compression (e.g. btrfs) and online block-level dedup (e.g. dm-vdo), that can aid in minimizing and/or avoiding actual block copies.
When all else fails, cp will anyway fallback to buffered IO, so it won't be bound by the rate of IO (as long as writeback can keep up with memory pressure).
Of course improvements on all layers are always welcome and there are certainly workloads that can benefit from an optimized raw throughput, low-latency and scalability. Demanding workloads are typically using dedicated IO stacks and bypass the kernel altogether (e.g SPDK), but the amazing efforts of Axboe and everyone else in the kernel community have been continuously bringing linux up to par in terms of IO performance.
[1] https://man7.org/linux/man-pages/man2/copy_file_range.2.html [2] https://git.savannah.gnu.org/cgit/coreutils.git/commit/?id=4...