> To make things worse, non-blocking I/O is done completely differently
> under Unix and under Win32. I'm not even sure Win32 provides enough
> support for async I/O to write a real user-level scheduler.
sigh, VMS got the link between processes, threads, I/O and waitable events (specifically, the link between tying the completion of future I/O to subsequent computation) right from day one. And by virtue of Cutler, therefore, so did NT, and thus, Windows.UNIX did not. The core concept of separating the work (computation to be done after an event occurs) from the worker[1] (the thread that performs the work) is absent; the manifestation of that is the lack of good, completion-oriented asynchronous I/O primitives. Instead of being able to say to the kernel "here, do this, then let me know when you're done"[2] and moving on to the next piece of work in the queue, you have to do the elaborate non-blocking multiplex dance for socket I/O, palm file I/O off onto a separate set of threads that can block (or do AIO) and generally manage all threading and concurrency primitives yourself.
It took me ten years of UNIX systems programming to suddenly grasp the elegance of the VMS/NT/Windows approach a few years ago. It provides you with everything you need to optimally exploit all your cores for work that is both heavily compute bound and I/O bound.
It has been fascinating to see the difference in performance between Linux and Windows in practice with PyParallel when Windows kernel primitives are exploited properly:
https://speakerdeck.com/trent/pyparallel-pycon-2015-language....
And more recently, with 10Gbe hardware at home:
Linux lwan (the top performer on Techempower Framework Benchmark):
[trent@zebra/ttypts/1(~s/wrk)%] time ./wrk --timeout 120 --latency -c 256 -t 12 -d 30 http://10.0.0.2:8080/plaintext
Running 30s test @ http://10.0.0.2:8080/plaintext
12 threads and 256 connections
Thread Stats Avg Stdev Max +/- Stdev
Latency 5.34ms 7.46ms 197.13ms 82.40%
Req/Sec 14.41k 364.49 18.82k 76.61%
Latency Distribution
50% 398.00us
75% 9.01ms
90% 17.50ms
99% 28.03ms
5178617 requests in 30.10s, 0.93GB read
Requests/sec: 172048.49
Transfer/sec: 31.67MB
Windows PyParallel: [trent@zebra/ttypts/1(~s/wrk)%] time ./wrk --timeout 120 --latency -c 256 -t 12 -d 30 http://10.0.0.2:8080/plaintext
Running 30s test @ http://10.0.0.2:8080/plaintext
12 threads and 256 connections
Thread Stats Avg Stdev Max +/- Stdev
Latency 1.52ms 9.38ms 492.43ms 99.33%
Req/Sec 18.37k 1.01k 22.75k 73.50%
Latency Distribution
50% 1.09ms
75% 1.28ms
90% 1.56ms
99% 5.18ms
6598900 requests in 30.10s, 1.03GB read
Requests/sec: 219236.69
Transfer/sec: 34.92MB
./wrk --timeout 120 --latency -c 256 -t 12 -d 30 106.30s user 138.87s system 814% cpu 30.114 total
[1]: https://speakerdeck.com/trent/parallelism-and-concurrency-wi...[2]: https://speakerdeck.com/trent/pyparallel-how-we-removed-the-...