IOCP merely allows IO completion to happen on worker threads in a threadpool rather than the particular thread that initiated the IO request. Many concepts in the NT kernel come from a long line of production kernels in mainstream DEC operating systems going back to RSX-11M and VMS. That entire lineage of operating systems (culminating in NT) all have true asynchronous IO using IO Request Packets (IRPs) within the kernel to initiate/queue IO requests and immediately return. There are individual cases within NTFS where blocking will occur but those are special cases rather than the rule as the kernel IO system in general is entirely async.
A thread pool was also more reliable than the kernel APIs at ensuring overlapped I/O for throughout. libaio (kernel API) calls sometimes turned into blocking calls depending on filesystem internals, in a way that thread pools don't.
POSIX AIO, on Linux, just uses a userspace thread pool. It's implemented in libc and is not particularly fast. You can do better with your own thread pool.
Perhaps Linux threads got faster, faster than the APIs got better.
Even with io_uring, in my tests on fast NVMe storage RAIDs where it should make the most difference, I found io_uring wasn't noticably faster than a well-optimised (for Linux) thread pool.
io_uring can potentally adapt better to shared workload and different cache situations, because it has access to kernel information that userspace thread pools are not granted. But I was surprised to find no significant random-access I/O performance increase between my thread pool and io_uring, when I was trying to optimise both methods to get the best throughput out of a system.
Tokio's filesystem handling in Rust is the same way by default - a pool of IO threads.
Edit: I just remembered that Avi Kivity (of KVM and ScyllaDB fame) wrote a tool to detect this: https://github.com/avikivity/fsqual
You need completion notifications instead (like IOCP or io_uring or POSIX aio). You could in principle start an async splice from an FD into a pipe and then poll for the pipe to be ready to read, but IIRC there is no async splice.