And no, it's not just because of io_uring it is faster. It's also because it's multi-threaded, has absolutely different hashtable design, uses a different memory allocator and many other reasons (i.e design decisions we took on the way).
we use io_uring for everything: network and disk. Each thread maintains its own polling loop that dispatches completions for I/O events, schedules fibers etc. Everything is done via io_uring API. All socket writes are done via ring buffer etc. If you run strace on DF you won't see almolst any system calls besides io_uring_enter