Benchmarking OS primitives
bitsnbites.eu
bitsnbites.eu
One of my major complaints with Windows is that things just 'feel slow'. I have to wait very often. Opening an FTP location? Wait for 5 seconds, (and it also opens in a new window, leaving the old window open in an unusable state - very confusing). Starting a GUI? Wait for 5 seconds.
My laptop and raspberry pi at home both work a lot smoother than the hi-end (it's a brand new Dell XPS machine with 8 GB RAM - which I consider hi-end) laptop I have at work.
I still find it hard to comprehend that people are buying ridiculously overpowered Windows computers for tasks like browsing and document editing. Developers are at fault too - if it runs smoothly on your $1000+ machine with 32GB RAM, that does not mean that the average user will be able to even use it. Everyone and their mother is jumping at the sustainability hype, but at the same time developers assume that everyone buys a new computer and phone every other year, for the same tasks we've been doing for decades. Once you realize this it's hard to use a Windows system and not cringe at the mess of laggy/unresponsive GUI's.
You're holding it wrong.
> My laptop and raspberry pi at home both work a lot smoother than the hi-end (it's a brand new Dell XPS machine with 8 GB RAM - which I consider hi-end) laptop I have at work.
If your RPi is faster at comparable tasks than that Windows PC, your Windows PC has some extremely serious setup problems.
(if your employer uses anything like the commercial security software mine does, that's one potential problem)
On Linux, I have several terminals open, which is just so much more lightweight.
Of course, from a performance point-of-view, these are not comparable, but that is exactly my point. On windows, everything has a GUI, and everything seems to assume a much more hi-end machine.
P.S. I admit that it was a bit misleading to post this rant under this article, since the article measures raw OS performance, while my point is that lower performance is usually sufficient too, if you use tools that have a single purpose instead of trying to be an OS in itself.
But the performance between the RPi and the Windows machine shouldn't be comparable at all, and if the RPi is coming out ahead for anything more trivial than opening a terminal window, I'd look at the setup of the Windows machine. There is plenty that can get effed up there.
OK, but .. how exactly? It's going to be Windows Defender, isn't it?
The oddball feature that causes Windows disk I/O to basically lock up on occasion, and I'm amazed that they haven't turned it off by default by now, is volume shadow copy and the automatic creation of restore points. It's bad on an SSD because it's very slow and uses a lot of space, and on a hard disk it's so slow that it ought to be criminal.
For example, time from boot to allow do something like browsing a web page. I don't did a precise measurement of time, but on Windows 10 I need to wait like 10 fucking minutes to allow to do something! And the hard disk doing a lot of horrible noises, so Windows must doing something. On linux, would take like a single minute or two, and I don't noticed any noticeable hard disk activity.
I don't know that is messed with Windows (probably I messed something on Windows), but I really hate this.
However, this isn't exactly new information. The general slowness of the OSX kernel has been known for years, via other benchmarks like lmbench. Its one of the reasons they were the first to implement a vdso-like interface for things like gettimeofday().
Benchmarking typical environments vs artificially lean is much more helpful in practice.
Most machines were stock configured (Ubuntu ext4 install for Linux, Windows 10 w/ Windows Defender, macOS X stock install, etc), so I believe that they should be representative for average users.
The benchmark suite is open source and easy to run on your own hardware if you like to get more accurate/representative figures for a particular setup.
However, poor or bad applications also play a role here. For some reason explorer.exe requires 1-2 orders of magnitude more I/O time than dir.exe, and somehow the local search is slower than a human binary search. However they managed to do that bad of a job remains a mystery.
find /cygdrive/c/ -type f -iname '*somefile.tar'
Maybe on proper UNIX (and BSDs) where you have the whole base system as more or less one package, but on Linux certainly not. Honestly the whole concept of "OS" does not make that much sense on traditional Linux distros. You could consider a whole distro an OS, but then the phrase is so all-encompassing that it becomes meaningless. On the other hand calling just the kernel "OS" is not quite right either.
> b) malloc has to get the memory from the OS eventually, via sbrk or mmap.
Yes, eventually and occasionally. But there is fairly big disconnect between malloc and OS. You got me curious, so I did run the code from the article with few different NUM_ALLOCS values to see how it behaves. These are the results from glibc malloc on my system:
Benchmark: Allocate/free 10000000 memory chunks (4-128 bytes)...
70.310378 ns / alloc
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
100.00 0.023622 4 5996 brk
Benchmark: Allocate/free 1000000 memory chunks (4-128 bytes)...
171.745062 ns / alloc
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
100.00 0.250319 17 15004 brk
Benchmark: Allocate/free 100000 memory chunks (4-128 bytes)...
33.700466 ns / alloc
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
100.00 0.000203 3 63 brk
Benchmark: Allocate/free 10000 memory chunks (4-128 bytes)...
30.589104 ns / alloc
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
0.00 0.000000 0 9 brk
Benchmark: Allocate/free 1000 memory chunks (4-128 bytes)...
26.941299 ns / alloc
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
100.00 0.000008 2 4 brk
Note how the syscall count blows up at NUM_ALLOCS=1000000, which just happens to be the original value from the article. Yes, I checked and glibc did not fall back to mmap in any of these cases. You can already start seeing why this might not be the best of benchmarks.Then just for fun, I tried using jemalloc, which is a drop-in replacement for standard malloc. These are the results:
Benchmark: Allocate/free 10000000 memory chunks (4-128 bytes)...
68.486404 ns / alloc
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
100.00 0.000186 4 47 mmap
0.00 0.000000 0 2 brk
Benchmark: Allocate/free 1000000 memory chunks (4-128 bytes)...
59.097052 ns / alloc
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
96.45 0.000544 16 33 mmap
3.55 0.000020 10 2 brk
Benchmark: Allocate/free 100000 memory chunks (4-128 bytes)...
54.659843 ns / alloc
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
0.00 0.000000 0 24 mmap
0.00 0.000000 0 2 brk
Well, well, well. It certainly paints a very different picture.In artificial tests like these, you frequently get the best performance by flushing data out as fast as possible, while in most "real-world" scenarios you have some temporal locality that makes keeping data around a win. Optimizing for these sorts of benchmarks can actually harm performance.
Still, fun.
http://blog.zorinaq.com/i-contribute-to-the-windows-kernel-w...
Missing are the results for the BSDs. I'm particularly interested on Dragonfly BSD. Maybe I'll try them myself when 5.2 is out, which will be soon.
Otherwise, the filesystem bench is pointless without:
- SSDs type. Tlc? SLC? Are they same, different?
- Linux filesystem type and fstab flags.
Linux filesystem: stock ext4
CreateProcess() is like posix_spawn(), or if you prefer fork()/exec().
Windows is a thread based OS, not process based, hence why the focus on thread performance, not on process creation.
Which, somewhat ironically, leads NT to have worse numbers in the create thread test than linux in the create process one (25.6us vs 18us).
The redeeming factor of NT is their async IO model which afaik is the best among mainstream OS.
The key difference is that I/O completion ports can be used to achieve asynchronous I/O on any underlying object, e.g. files and sockets, and they have this nifty built-in concept of concurrency, such that the kernel can ensure there is always one running thread per CPU core (which is optimal from a scheduling perspective).
You can't use file descriptors with epoll/kqueue, and you certainly can't say "ensure every core only has one active thread running".
"The key to understanding what makes asynchronous I/O in Windows special is...": https://speakerdeck.com/trent/pyparallel-how-we-removed-the-...
"Thread-agnostic I/O with IOCP": https://speakerdeck.com/trent/pyparallel-how-we-removed-the-...
The concepts of processes and threads work just the same in Linux and Windows (and internally just map to the execution unit of the scheduler, together with resource mappings and privileges), and user-space expectations are similar for the two. The main difference is that fork() is not available on Windows, but fork() is a terrible idea anyway.
Fast spawn of processes isn't used for performance critical things on either OS, as process spawning is considered slow on Linux and entirely useless on Windows. Fast spawn of threads is also generally avoided, as even that is usually considered too slow.
Windows is slow at creating processes (and most other things involving the kernel) not because of differences in OS use-case, but simply due to performance apparently not being a priority for Microsoft.
A thread based OS is an OS where threads are the core unit of execution, and processes are just a kind of execution capsule with one thread executing by default.
The kernel scheduler only understands threads.
This by opposition to process based OSes like UNIX, where there is a clear distinction between a process and thread execution.
The kernel scheduler handles processes and threads separately.
In many UNIX platforms, a process that doesn't perform any thread related API call, won't have any thread running on its context.
This was quite clear during the days when UNIX systems where still researching how to adapt threads into the process execution model.
And in many cases the impedance mismatch is still visible in modern UNIX systems, like for example what happens to any given thread when a signal is triggered, or to the whole process when a thread decides to fork.
You can start by getting yourself a copy of "Windows Internals" book.
Here is an old version of "Processes, Threads, and Jobs in the Windows" chapter in the 5th edition.
https://www.microsoftpressstore.com/articles/printerfriendly...
However, the topic would appear to be Windows, Linux and potentially also macOS. That's what the benchmarks are about. No one mentioned other OS's.
This only really matters when you're trying to understand how PID, PPID, and TGID fit together and why there is no TID, though.
To quote the FreeBSD manual: "Traditional UNIX® does not define any API nor implementation for threading, while POSIX® defines its threading API but the implementation is undefined."
macOS was at least temporarily considered a true UNIX, and the scheduling primitive there is a Mach task. I frankly don't remember much about FreeBSD anymore, but I would assume that the unit of scheduling there is somewhat identical to that of Linux... Just implemented nicer.
The primary difference between Windows and Linux (which is the topic at hand, not other Unixes or esoteric OS's) is in terminology. A Windows "process" is not the same as a Linux "process". A Windows "thread" is not the same as a Linux "thread".
However, a Windows "execution resource" ("thread") is quite identical to a Linux "execution resource" ("process"), or the macOS "execution resource" ("Mach task", not "process" or "thread").
Windows and Linux also have "execution resource groups", in the form of "process" and "thread group"/"parent process", respectively. They are implemented slightly differently (dedicated device vs. "master" execution resource), but the end-result is similar.
These constructs implement identical functionality for all intents and purposes (the differences are just in some minor limitations and API choices). The scheduler only operates on the execution resource, but might look at the execution resource group when making scheduling decisions. This is shared between all the OS's, and is a minor implementation detail that can change between releases.
The distinction between "thread-based" and "process-based" does not exist. Linux is a heck of a lot faster to create "resource groups" than Windows is, but that is due to better code, not fundamental design limitations.
(Of course, an esoteric OS might implement something entirely different from the concepts of "processes" and "threads", but that's a fun discussion for another day—the important thing is the contemporary OS's are all the same.)