CreateProcess() is like posix_spawn(), or if you prefer fork()/exec().
Windows is a thread based OS, not process based, hence why the focus on thread performance, not on process creation.
Which, somewhat ironically, leads NT to have worse numbers in the create thread test than linux in the create process one (25.6us vs 18us).
The redeeming factor of NT is their async IO model which afaik is the best among mainstream OS.
The key difference is that I/O completion ports can be used to achieve asynchronous I/O on any underlying object, e.g. files and sockets, and they have this nifty built-in concept of concurrency, such that the kernel can ensure there is always one running thread per CPU core (which is optimal from a scheduling perspective).
You can't use file descriptors with epoll/kqueue, and you certainly can't say "ensure every core only has one active thread running".
"The key to understanding what makes asynchronous I/O in Windows special is...": https://speakerdeck.com/trent/pyparallel-how-we-removed-the-...
"Thread-agnostic I/O with IOCP": https://speakerdeck.com/trent/pyparallel-how-we-removed-the-...
The concepts of processes and threads work just the same in Linux and Windows (and internally just map to the execution unit of the scheduler, together with resource mappings and privileges), and user-space expectations are similar for the two. The main difference is that fork() is not available on Windows, but fork() is a terrible idea anyway.
Fast spawn of processes isn't used for performance critical things on either OS, as process spawning is considered slow on Linux and entirely useless on Windows. Fast spawn of threads is also generally avoided, as even that is usually considered too slow.
Windows is slow at creating processes (and most other things involving the kernel) not because of differences in OS use-case, but simply due to performance apparently not being a priority for Microsoft.
A thread based OS is an OS where threads are the core unit of execution, and processes are just a kind of execution capsule with one thread executing by default.
The kernel scheduler only understands threads.
This by opposition to process based OSes like UNIX, where there is a clear distinction between a process and thread execution.
The kernel scheduler handles processes and threads separately.
In many UNIX platforms, a process that doesn't perform any thread related API call, won't have any thread running on its context.
This was quite clear during the days when UNIX systems where still researching how to adapt threads into the process execution model.
And in many cases the impedance mismatch is still visible in modern UNIX systems, like for example what happens to any given thread when a signal is triggered, or to the whole process when a thread decides to fork.
You can start by getting yourself a copy of "Windows Internals" book.
Here is an old version of "Processes, Threads, and Jobs in the Windows" chapter in the 5th edition.
https://www.microsoftpressstore.com/articles/printerfriendly...
However, the topic would appear to be Windows, Linux and potentially also macOS. That's what the benchmarks are about. No one mentioned other OS's.
This only really matters when you're trying to understand how PID, PPID, and TGID fit together and why there is no TID, though.
To quote the FreeBSD manual: "Traditional UNIX® does not define any API nor implementation for threading, while POSIX® defines its threading API but the implementation is undefined."
macOS was at least temporarily considered a true UNIX, and the scheduling primitive there is a Mach task. I frankly don't remember much about FreeBSD anymore, but I would assume that the unit of scheduling there is somewhat identical to that of Linux... Just implemented nicer.
The primary difference between Windows and Linux (which is the topic at hand, not other Unixes or esoteric OS's) is in terminology. A Windows "process" is not the same as a Linux "process". A Windows "thread" is not the same as a Linux "thread".
However, a Windows "execution resource" ("thread") is quite identical to a Linux "execution resource" ("process"), or the macOS "execution resource" ("Mach task", not "process" or "thread").
Windows and Linux also have "execution resource groups", in the form of "process" and "thread group"/"parent process", respectively. They are implemented slightly differently (dedicated device vs. "master" execution resource), but the end-result is similar.
These constructs implement identical functionality for all intents and purposes (the differences are just in some minor limitations and API choices). The scheduler only operates on the execution resource, but might look at the execution resource group when making scheduling decisions. This is shared between all the OS's, and is a minor implementation detail that can change between releases.
The distinction between "thread-based" and "process-based" does not exist. Linux is a heck of a lot faster to create "resource groups" than Windows is, but that is due to better code, not fundamental design limitations.
(Of course, an esoteric OS might implement something entirely different from the concepts of "processes" and "threads", but that's a fun discussion for another day—the important thing is the contemporary OS's are all the same.)