Assuming you have an hardware DMA ring buffer for I/O that is directly mapped in userspace, the only thing that is really needed is to write the operation type, size, disk position and memory position to the ring buffer, update the buffer position and check for flush, doable in around 8 CISC instructions (plus the slowpath), so around 100x inefficient.
Without the direct mapped ring buffer and with a filesystem, you need a kernel to translate from uring to the hardware ring buffer, and here it still seems around 10x inefficient as around 100 instructions should be enough to do the translation (assuming pages already mapped in the IOMMU, that you have the file block map in cache, and that the whole system is architected to maximize the efficiency of this operation).
This type of thing is not measured in instructions anymore, it's measured in off-board operations and latencies. Cache misses, latencies to poke MMIO and get back interrupts if necessary, DMA transfer to complete, device access time. In this case it seems the hardware is theoretically capable of about 12M so the core mostly be just waiting for that.
For some workloads it's not that hard to have that deep queues. What's harder is to know when to use them and when not. There's really not enough information available to make any of this self-tuning.
EDIT: 512, looking at the screenshot..