Edit: I think I found it [1] That iteration was from 2002. I'd be curious to see if his opinion has evolved in 12 years.
Also, interesting to see game developer Chris Hecker [2] in that thread.
[1] http://caml.inria.fr/pub/ml-archives/caml-list/2002/11/threa... search for "Why systhreads?" and "Xavier Leroy". Also, damn their website's broken.
Better link, I wish GMane had better Googlejuice: http://thread.gmane.org/gmane.comp.lang.caml.general/16381/f...
> To make things worse, non-blocking I/O is done completely differently
> under Unix and under Win32. I'm not even sure Win32 provides enough
> support for async I/O to write a real user-level scheduler.
sigh, VMS got the link between processes, threads, I/O and waitable events (specifically, the link between tying the completion of future I/O to subsequent computation) right from day one. And by virtue of Cutler, therefore, so did NT, and thus, Windows.UNIX did not. The core concept of separating the work (computation to be done after an event occurs) from the worker[1] (the thread that performs the work) is absent; the manifestation of that is the lack of good, completion-oriented asynchronous I/O primitives. Instead of being able to say to the kernel "here, do this, then let me know when you're done"[2] and moving on to the next piece of work in the queue, you have to do the elaborate non-blocking multiplex dance for socket I/O, palm file I/O off onto a separate set of threads that can block (or do AIO) and generally manage all threading and concurrency primitives yourself.
It took me ten years of UNIX systems programming to suddenly grasp the elegance of the VMS/NT/Windows approach a few years ago. It provides you with everything you need to optimally exploit all your cores for work that is both heavily compute bound and I/O bound.
It has been fascinating to see the difference in performance between Linux and Windows in practice with PyParallel when Windows kernel primitives are exploited properly:
https://speakerdeck.com/trent/pyparallel-pycon-2015-language....
And more recently, with 10Gbe hardware at home:
Linux lwan (the top performer on Techempower Framework Benchmark):
[trent@zebra/ttypts/1(~s/wrk)%] time ./wrk --timeout 120 --latency -c 256 -t 12 -d 30 http://10.0.0.2:8080/plaintext
Running 30s test @ http://10.0.0.2:8080/plaintext
12 threads and 256 connections
Thread Stats Avg Stdev Max +/- Stdev
Latency 5.34ms 7.46ms 197.13ms 82.40%
Req/Sec 14.41k 364.49 18.82k 76.61%
Latency Distribution
50% 398.00us
75% 9.01ms
90% 17.50ms
99% 28.03ms
5178617 requests in 30.10s, 0.93GB read
Requests/sec: 172048.49
Transfer/sec: 31.67MB
Windows PyParallel: [trent@zebra/ttypts/1(~s/wrk)%] time ./wrk --timeout 120 --latency -c 256 -t 12 -d 30 http://10.0.0.2:8080/plaintext
Running 30s test @ http://10.0.0.2:8080/plaintext
12 threads and 256 connections
Thread Stats Avg Stdev Max +/- Stdev
Latency 1.52ms 9.38ms 492.43ms 99.33%
Req/Sec 18.37k 1.01k 22.75k 73.50%
Latency Distribution
50% 1.09ms
75% 1.28ms
90% 1.56ms
99% 5.18ms
6598900 requests in 30.10s, 1.03GB read
Requests/sec: 219236.69
Transfer/sec: 34.92MB
./wrk --timeout 120 --latency -c 256 -t 12 -d 30 106.30s user 138.87s system 814% cpu 30.114 total
[1]: https://speakerdeck.com/trent/parallelism-and-concurrency-wi...[2]: https://speakerdeck.com/trent/pyparallel-how-we-removed-the-...
Digital dropped the ball in the late 80s with regards to management of Cutler and his team, canceling his PRISM project and leaving him and his team disgruntled.
Elsewhere in Seattle, a chap named Bill Gates was flush with billions of cash and knew that the shelf life of DOS was limited; if Microsoft were to succeed, they needed a new, robust, reliable and high-performance OS that they could "bet the company on".
Gates got word that Cutler was disgruntled at Digital, and a mutual party set up a meeting. Cutler was dismissive of Microsoft's technology stack at the time (DOS and some office apps) -- he was a hardcore OS engineer, and DOS was a toy.
Gates persisted, ensuring Cutler that he would have the opportunity to build the next generation of OS from the ground up and essentially unlimited resources at his disposal to do it. Cutler eventually agreed, and the NT kernel project was born.
http://www.amazon.com/Show-Stopper-Breakneck-Generation-Micr...
http://windowsitpro.com/windows-client/windows-nt-and-vms-re...
Reading the book and learning the story behind NT's development, it's just amazing that such a good OS came out of that process - they released years after their initial projections and were rushed the whole time. But of course the really good parts of NT - the kernel, the object manager, the pager, async IO, the threading model - were things Cutler and his cohorts had been working on for years, first with VMS, then with PRISM, and then finally in NT. They had YEARS to ruminate about those things before they ever arrived at Microsoft.
The bits of NT that aren't so well-regarded - the registry, NTFS, the graphical shell, csrss.exe and the 'microkernel' design - were completely new and developed in much less time and with less practical experience behind them than they really deserved.
I sent a tweet to the author saying I was really enjoying the book when I was about half way through it and he actually e-mailed me to say thanks. How nice is that!
Arguably, Linus' greatest work was Git, not Linux. Linux is, architecturally, a piece of shit! Actually, wait, so is Git. Mercurial does everything Git does and does it far better and more elegantly. So yeah, wait... one wonders where Linus gets all his fanatics from!
[Edit: Clarification]
This doesn't mean that Windows' philosophy does not give you optimal performance in PyParallel. It simply means that OCaml had chosen for its low-level system primitives a Unix model and that it was difficult to make a Windows version of the same primitives so that OCaml programmers could write this kind of program portably between Windows and Unix.
NOTE: without, at the time it is in my timezone, looking up the full post, I have to say that I don't think that the quoted two sentences have anything to do with the discussion. It seems to me that the two sentences assume that a multicore (multiprocessor, at the time the post was written) OCaml runtime is not available, and discusses the options to still provide threads. A user-level scheduler is one option to provide threads to OCaml programs without a concurrent OCaml runtime. Another option is to use Windows' native threads and superior philosophy for blocking primitives to run each OCaml thread as a native thread (although at most one of these will be running at any given time. All the others will be waiting on the heap mutex).
OCaml ended up providing threads under Windows and a Unix-like “Unix” module around 1996-ish, way before the linked discussion. So thanks for the explanation about VMS, but I think it is off-topic, too.
NOTE 2: I have now read the original post. You should, too. It starts with:
> Threads have at least three different purposes:
>
> 1- Parallelism on shared-memory multiprocessors.
> 2- Overlapping I/O and computation (while a thread is blocked on a network
> read, other threads may proceed).
>3- Supporting the "coroutine" programming style
> (e.g. if a program has a GUI but performs long computations,
> using threads is a nicer way to structure the program than
> trying to wrap the long computation around the GUI event loop).
>
> The goals of OCaml threads are (2) and (3) but not (1) (for reasons
> that I'll get into later)
What makes it relevant to the current discussion is (1), but Xavier is discussing (2) and (3) at the time of the quote you chose to take out of context.
I'm not disputing any of the technical things he's saying; just ranting about the unfortunate nature of two vastly different kernel models, and the fact that no open source stuff properly exploits Windows facilities, despite them being technically superior.
It'd be interesting if running under the VMS/NT thread/fork model could be seen as a reason to deploy some apps on ReactOS rather than Linux/BSD. Would also be interesting if one could see any difference running a multi-core KVM guest on ReactOS vs a Linux/BSD guest/container/jail. Although I suppose one would need to dedicate a hw nic to see any real results (avoiding the host OS/VM scheduler etc)?
Note-to-self: something to play with...
I was also curious to see what would happen if I tried to install it on Wine.
> It'd be interesting if running under the VMS/NT thread/fork model could be seen as a reason to deploy some apps on ReactOS rather than Linux/BSD.
I... couldn't imagine trying to use ReactOS instead of Windows for an actual deployment of anything. Why wouldn't you just use Windows? (Serious question.)
This isn't academic -- look at Sun OS/Solaris. Granted we have open Solaris etc... but that appears as an accident of timing more than anything -- in retrospect.
Now, for the more relevant part: ReactOS vs Windows: If all you want is the kernel/thread model I could see going with ReactOS (pending actual research, as in: does it actually work :-). If you're deploying SQL Server/IIS .net (pending the so far seemingly serious effort to open .net) -- I don't know why one wouldn't go with Windows, no. In that scenario you'd be beholden (good and bad) to Redmond either way.
But for something like a python fork -- I could see something like ReactOS (or any other alternate kernel) be an interesting thing. You don't need much from the OS -- just classic services: basic filesystem/persistence, perhaps privilege separation (not so important for micro-service vms), scheduling.
> What about hyperthreading? Well, I believe it's the last convulsive movement of SMP's corpse :-)
Oh how things have changed. This was written before it was clear just how much of a disaster the P4 was, so it was a pretty reasonable position at the time.
"In summary: there is no SMP support in OCaml, and it is very very unlikely that there will ever be. If you're into parallelism, better investigate message-passing interfaces."