An EEVDF CPU Scheduler for Linux
lwn.net
lwn.net
If I’m running a video encoding using all available CPUs, I don’t want that to cause lag with all my other processes.
This happens to me a lot when running node processes like unit tests that running on all cores. It slows everything else down. I even set the nice value on node to always be the lowest priority and it only helps a little.
I’ve actually had Linux become completely unresponsive when running many large node instances. That shouldn’t happen.
I'm not clear on what mechanisms exist today for preventing things like a DoS of GPU resources in something of a single entry to the GPU driver without returning to userspace. We've all seen things like GPU stalls reported in dmesg, so clearly the drivers are tasked with preventing that sort of hang where a process enters the driver and stays there far too long.
"man 7 sched" provides a lot of detail about the real-time scheduling policies and options.
(Of course if the compositor is essential to the kinds of interaction you need, e.g via the GUI, you'll still be stuck if the compositor stops working. But that's not a real-time priority problem.)
A couple more things to try:
Reduce swappiness. Linux tries to swap out idle pages to free memory for block cache, even when there's only light memory pressure. If you have a dozen processes working hard and your UI is untouched for several minutes, it can get swapped out and lag.
sudo sysctl vm.swappiness=1
Set the CPU affinity for your build processes to exclude CPU 0, ensuring one is always free for UI: taskset 0xFFFFFFFE nice your_build_command_here
Neither is perfect but they can help.[1] may help if you run such a workload from CLI.
It also boosts the priority of threads that get woken up after waiting for IO, events, etc, as opposed to being CPU bound. Which I guess Linux's fair scheduler also kind of ends up doing, even if via an entirely different mechanism (simply by observing that they used less CPU).
It's interesting how different operating systems take completely different approaches to their schedulers. Linux seems to try to make quite sophisticated schedulers, trying out very different concepts, but keeps the scheduler very seperated and "blind" to the rest of the system. Meanwhile Windows has an incredibly simple scheduler (run the highest-priority runnable thread, interrupt it if its timeslot is over or a more important thread become ready, repeat) and puts all the effort into letting the rest of the kernel nudge priorities and timeslot lengths based on what the process or thread is doing.
Of course if you really want to know, the source code of Windows XP got leaked and is easy enough to find on github, so it should be possible to verify. I don't really know my way around that ball of source code though.
1: https://www.microsoftpressstore.com/articles/article.aspx?p=...
It's so light, even when a system is highly loaded, it still works fine (in my experience anyway)
Well I've seen it but it tended to be RAM/memory rather than CPU originated.
Thanks for the update!
Run the encoder at a lower priority?
This is almost certainly not a scheduling problem but a swapping problem. Linux is famously bad at handling these kinds of situations.
About five years ago I decided to never use swap anymore, and I couldn't be happier with that decision.
How does that make sense? Either your workload fits into RAM, so no swapping occurs and everything is responsive. Or your workload is too large to fit into RAM so
a) you have a swap partition and you system starts swapping and everything becomes really unresponsive, but the task will finish.
b) you don't have swap so the kernel kills the process using most memory, which is most likely the thing you're just trying to run.
How can you be happy with option b? I mean obviously get more RAM either way, but I'd rather have that safety net for the occasional memory intensive task. Obviously not a solution if you're short on RAM by a huge margin but if it's just a GB or two it's acceptable, because it's likely that the kernel can just swap out stuff from background programs that don't actively access their stuff right now.
You can see this well from the kennel's proactive swapping. If I run a bog standard desktop Linux distro and leave swappiness at the default 60, after a few hours of normal usage where I don't even get close to filling up my RAM with actual memory allocations from processes, the kernel will have swapped out almost a GB of data from several background services. Because a lot of programs allocate some memory for something and then never look at it in a long time. So it makes more sense to swap that to disk and instead use that ram as page cache.
Also, with current NVMe speeds, swapping has become a lot more bearable. Really the only place where I can see why you wouldn't use swap is servers with a well defined, dedicated workload. Here I indeed want a fast and noticable fail mode in case a new version of whatever I'm running might have a memory leak, or just vastly different behavior regarding memory usage.
When a workload runs out of RAM, it often really runs out of RAM, and would just be swapping continuously. It may even run out of swap space, if it was left to run for long enough. In my experience, even with NVMe drives, it just wasn't worth it because it usually wouldn't be able to finish in a reasonable amount of time. In part, I'm sure that's because the Linux kernel behaved much worse than is theoretically possible.
This may obviously depend on your workload, but I've heard similar sentiments from others (in fact, it was somebody else who suggested disabling swap to me based on their experience). The other point is that a sibling comment mentions some recent kernel changes that may have led to improvements, but I haven't given those a chance.
> The "Earliest Eligible Virtual Deadline First" (EEVDF) scheduling algorithm is not new; it was described in [a] 1995 paper by Ion Stoica and Hussein Abdel-Wahab. Its name suggests something similar to the Earliest Deadline First algorithm used by the kernel's deadline scheduler but, unlike that scheduler, EEVDF is not a realtime scheduler, so it works in different ways.
On a more controversial note, I trust more Meta & Google engineers proposing alternative schedulers that have been properly AB tested on millions of servers running a wide range of heterogeneous software VS small-scale synthetic benchmarks even when run by experts/maintainers.
[1]: https://lwn.net/ml/linux-kernel/20221130082313.3241517-1-tj@...
This is why I greatly desire the death of Linux. Its major subsystems are maintained by people with the social skills of a hyena and little to no relevant field experience.