Buffer allocation with IOCP feels incredibly ugly. You must allocate a buffer before you initiate a read(), and then you must leave that buffer alone until the read() completes. You would think that this has the advantage that the kernel doesn't need to allocate its own buffer, and you get some sort of zero-copy magic where bytes land directly in userspace. But that's not really the case. The socket still needs a buffer on the kernel side in case bytes arrive while no read() is pending. So now you have to allocate a buffer for each socket and the kernel also has to allocate a buffer for each socket, which seems like a waste.
This rabbit hole gets deeper. What happens if the socket receives two packets in rapid succession? In a naive implementation, the first packet completes the read(), and signals the completion on the IOCP. The app now has to process that event and start a new read() when it's ready. But in the meantime, the second packet arrives. Whoops, no read() is pending, so now the packet has to go into a kernel-side buffer only to be copied later.
But that's a naive implementation. Apparently, in reality, the kernel implements some sort of nagle-like algorithm where it tries to wait a bit for additional packets before it actually signals completion of a read(). But this introduces delay, and many projects (e.g. Chrome) have discovered this delay is rather harmful to certain kinds of performance. I read somewhere that Chrome and others have given up on IOCP and use WSAPoll instead -- but I can't remember where I read this because all the lore about Windows event handling is hidden in random forum threads and Github gists rather than proper documentation.
It seems to me that the right way to do what Windows was trying to do here would be for userspace to allocate a ring buffer for the kernel to use, and then the app and the kernel would coordinate the start and end pointers of the ring buffer. Then the kernel needs no buffer of its own; it can always deliver to userspace. If the buffer fills up, the kernel can do exactly what it would do in the case of a regular kernel-side buffer filling up -- apply backpressure and force the peer to retransmit later.
On Linux, apparently, the kernel does not actually allocate static socket buffers. Instead, the "socket buffer" is actually something like an array of pointers to packets. When the read buffer is empty, it's not taking any space.
So on Linux, an idling socket is very cheap (by my understanding, at least).
The Windows kernel might do something similar internally (I don't know), but there's no way to avoid the redundant buffer on the userspace side when using IOCP.
So basically, it seems that IOCP turns out to be a rather-poor interface for networking in practice, despite at first appearing theoretically superior.
In general I think that readiness notification is superior to completion notification for network reads. Completion notification works fine for network writes and it is superior for disk IO (where readiness is ill defined).
The biggest advantage of IOCP though is that it basically acts as a dynamically sized threadpool, as it keeps the amount of running threads to a minimum. MacOS does something similar with GCD. I still haven't found a non-hacky way to do the same in Linux (you can sort-of emulate it with either with sched_fifo realtime threads or by using sched_idle to detect idelness).