The Linux Scheduler: A Decade of Wasted Cores (2016)
blog.acolyer.org
blog.acolyer.org
If I recall correctly, 2000 is 2.3.x which had a braindead trivial scheduler. Basically it just looped through the process list and executed goodness() and found the 'best'. This was obscenely slow when there were a lot of dormant processes. The process table links all mapped to the same cache set which led to cache evictions during the scheduling loop, even TLB evictions. Et cetera. It was just bad, really really bad. Compared to its BSD, Solaris, ... contemporaries, it was garbage.
The O(1) scheduler starts in 2.4 at the tail end of 2000 followed by the Completely Fair Scheduler in 2007, etc. And the Linux scheduler continues to get better. But in 2000, it sucked. Reeked.
This makes the overall scheduling problem much harder, to the point that they were built with a special "Disable half the cores" mode and supporting hardware to give both memory banks same-speed access to the remaining ones.
I own one and in both work and play I have had zero issues. If I drop a frame here and there in a game due to some memory latency? Eh, could care less. If you can afford a Threadripper you can afford a 1080 Ti and a Gsync monitor to smooth out any issues you might run into.
It would be more difficult to do it automatically, but if NUMA systems become more common then I see no reason why it shouldn't be tried.
Which means in reality, you could name approximately everyone that ran into this issue on a single list: top500.org.
The Linpack benchmark is an example of an HPC code that should run one process per core and be pinned.
We didn't have HPC workloads, just Postgres, which uses one OS process per connection, and performance was terrible as a result.
I'd bet, but not too much, that that was more due to a) postgres' internal locking implementation scaling horribly at that time b) zone_reclaim_mode leading to bad behaviour around IO.
It's really hard to find a decent article on the subject, here's two of the best I've found. https://www.mssqltips.com/sqlservertip/4403/understanding-sq... https://docs.microsoft.com/en-us/sql/relational-databases/th...
The better question is: why has the Linux Foundation been silent about it?
If this were true -- "a decade of wasted cores", with losses of "13-24% for typical Linux workloads" (for a decade, as the title suggests) -- then companies like Netflix would have lost many millions due to our choice of Linux, and due to the Linux scheduling maintainers and community failing to identify such egregious problems. The industry as a whole -- including every device and server that runs Linux -- would have lost many BILLIONS. It would be one of the most costly failures in technology EVER.
And the Linux Foundation is silent about this? Seriously?
Then why make a cryptic comment like you did? Cryptic comments kind of invite questions.
There's some extra research I did that I haven't shared yet, including, for example, when the bugs were introduced (still needs double checking):
- Bug 1: Mar 2011, for Linux 2.6.38
- Bug 2: Apr 2012, for Linux 3.4
- Bug 3: Dec 2009, for Linux 2.6.32
- Bug 4: Feb 2015, for Linux 3.19
This paper was published in early 2016, with the title "A decade of wasted cores".https://news.ycombinator.com/item?id=11573375 https://news.ycombinator.com/item?id=11502221
https://technet.microsoft.com/en-us/library/aa175393(v=sql.8...
The Windows Database team has been using the "User Mode Scheduling" feature of Windows to implement custom scheduling for databases for some time.
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
fair.c deadline.c rt.c