As far as I can tell, everything involving sockets goes through tons of queues like the ones he describes; concurrency, not parallelism.
If I kick off reads on ten different sockets, when the read completes the kernel has to put the thread waiting for the read onto a 'ready to run' queue before any work can actually happen. Scale up to 1000 pending reads and now that queue could get pretty full pretty quick, and lots of data is moving in and out. As I understand things, select and epoll improve on this by making the queue more efficient: now instead of using a thread per socket, and having the OS queue up threads on the ready-to-run queue, the OS helps you manage a queue of 'ready sockets', and you respond to those manually. But there's still a queue!
Are there IO APIs out there in modern OSes that let you reduce queueing in such a way? Like, hypothetically if performing a read on a socket required me to pass in a callback that would process the data from the read, the OS could respond to incoming data on that socket by immediately running the callback, whether it's on an OS-managed user-mode thread pool or via some other mechanism. This would at least reduce the amount of queueing going on - at most you have a queue of tasks for a thread pool that are responding to complete IO (similar to epoll, I guess?) but I bet there are ways to improve on that too.
(If memory serves, IO Completion Ports on NT have functionality that looks like this... but I bet it still involves lots of queueing and scheduling)