Assuming you have an hardware DMA ring buffer for I/O that is directly mapped in userspace, the only thing that is really needed is to write the operation type, size, disk position and memory position to the ring buffer, update the buffer position and check for flush, doable in around 8 CISC instructions (plus the slowpath), so around 100x inefficient.
Without the direct mapped ring buffer and with a filesystem, you need a kernel to translate from uring to the hardware ring buffer, and here it still seems around 10x inefficient as around 100 instructions should be enough to do the translation (assuming pages already mapped in the IOMMU, that you have the file block map in cache, and that the whole system is architected to maximize the efficiency of this operation).