It's a heck of a problem to deal with, and a OOM killer will go part of the way to fixing it, but I still have two complaints about the thrash problem.
1. In similar situations on other operating systems, I don't usually see the GUI freeze to an unusable state. In Linux hitting a thrash situation will freeze the computer to the point where Ctrl-Alt-F2 will not switch the vterm even when waiting several minutes, and it has to be rebooted. I suppose it's possible that I just haven't used other OSes enough in a long time and they have this too, but this seems like a solvable problem. I assume it's not a CPU issue, since the scheduler should handle that just fine. So why not have a mechanism by which the user's desktop environment and init / login system can reserve itself enough memory to always be responsive under any memory condition?
2. Browsers. The state of browser memory management is atrocious. Just about every time I've run out of memory (other than compiling some complex software with -O3) it's been because my browser is hogging a huge chunk of it (even if the OOM killer blames something else). Now, I understand that in theory unused memory is wasted memory. So a browser using a ton of memory can be a good thing. But this is only true when memory the browser is holding but not using can be reclaimed to be used by another program, and this seems to basically never happen. If I close and reopen the browser (with the tabs automatically restored) the memory usage drops by 80% or more; the issue is that the browser simply never suspends tabs I haven't used in days or weeks. Modern browsers are not good "desktop citizens", if you will. They hog memory to the point where opening anything new will create a pressure stall.
I suspect solving these two issues would largely fix thrashing for most desktop users, without the need to ever kill anything, which is obviously undesirable.
Feel free to tell me if one or both of these is incorrect or impossible. These are just my observations, I'm not a kernel dev or anything close to one.
You can recover system from this state by manually triggering OOM killer using Magic SysRq keys. (Alt+SysRq+F, provided you have SysRq enabled in /proc/sys/kernel/sysrq)
Because it's not the memory shortage that's nailing the UI. It's disk contention. When pathological swap activity causes disk thrashing, every process that needs to use the disk is now waiting in a very long queue with all the swap page in/out requests and every other process that is now backed up.
If you have a metrics agent that emits a handful of metrics on a per second granularity, you can see a bunch of things happen:
* Memory rises
* Some pages get swapped out
* Even if the system has no other processes writing to disk, you'll see the swap-out correlating with disk writes
* Memory rises some more
* You see more pages get swapped out
* (...and see more disk writes)
* Memory rises some more
* You see pages get swapped out and in, because now we're shuffling out pages that are imminently required by another running process.
* Disk queue lengths rise
* One by one, all processes are waiting (driving up the CPU "wa" number, and causing load averages to shoot up)
This is why, if you're oncall and you're trying to remotely troubleshoot "high load", you likely can't log in because the system needs to access /etc/passwd, /etc/nsswitch.conf, /etc/ldap/ldap.conf, /bin/bash... Your login gets handed a ticket with "∞" written on it, and is told to stand in the line for the disk.
But aren't these files already in the disk cache? Possibly. But another thing that happens when the system runs out of memory is that disk caches get purged. So you have to go back to the disk again. Also, when disk caches get dropped, your disk read performance is dogshit.
So what can you do? Try zswap? Use a different disk devices for swap? Try to pin important processes so the can't be swapped out?
NOTE: I might not be right about this stuff - I've never worked on the Linux kernel, and I rarely even compile it these days. But I've worked as SRE / *Ops for years and saw this problem a lot in badly configured systems. Same patterns all the time.
I suppose the solution is to figure out how to reduce or eliminate the number of disk reads after the DE initially loads, or cache everything needed in memory with a "do not evict" flag on it. Especially for the TTY, that's basically a crucial system function and keeping everything needed for it always available in memory wouldn't take much space relative to a typical desktop system.
I think at least some of the problem is the program code pages getting evicted to disk, so they have to be read back to memory too, which goes to the back of the disk queue. When I'm experiencing thrashing, even basic stuff like alt-tab to switch windows doesn't work. And I assume the window manager doesn't need to do disk reads to switch windows. Even the damn X cursor freezes!
I was able to get glitch-free audio by tweaking software and soft-irq realtime and non-realtime priorities, when the default Ububtu (with low-latency kernel) was nearly unusable for audio.
Which problem? Yes, there's always disk contention. But if you don't have swap enabled then you won't have pathological swap activity dominating the disk.
There are several bad things that happen when you run out of memory. One of them is pathological swapping. Another is loss of disk caches. Another is (eventual) oom. The one that has the most detrimental effect, and typically results in the system having to be reset, is swap activity saturating disk io. Even when you've got fast SSDs, because they're still orders of magnitude slower than memory.
Thrashing and becoming unusable for a long time.
This! Always wanted a /proc/$PID flag to flag process' pages never to go to swap and reserve a number of cache pages specifically for the process. Something similar to mlock.
Perhaps this falls outside the category of "any memory condition"... :)
Though, it was responsive. :D
Needless to say, I think 1 is harder than expected. Fedora has made some recent changes on the OOM killer front though, perhaps those are closer to what you're interested in.
* I run separate cgroups for admin SSH with (1G min, 4G max ) - worst case scenario I can ssh into the desktop and kill offending processes.
* Separate cgroup for two browsers ( 8G each max ).
* Separate cgroup per qemu VM ( VM max ram + 512M each max ).
After implementing that I went from periodic lock ups to a quick browser crash and a restart maybe once a week.
The "Auto Tab Discard" extension[1] for Firefox unloads tab contents, keeping only the title and favicon. Tabs can be put in an allowlist to prevent discarding. I don't know if Chrome or Safari have something similar.
[1] https://addons.mozilla.org/en-US/firefox/addon/auto-tab-disc...
> Modern browsers are not good "desktop citizens", if you will. They hog memory to the point where opening anything new will create a pressure stall.
I went the ballistic route: most of my browsers run inside local Docker containers, on which I put CPU and RAM quota.
It's my machine: I decide how "citizens" behave : )
#!/bin/sh
systemd-run \
--user \
--scope \
--property=MemoryHigh=4G \
--slice=background \
/usr/bin/firefox "$@"
Theoretically this does not prevent Firefox from using more than 4 GBs of memory (you need MemoryMax for that), but practically it does. MemoryHigh=bytes
Specify the throttling limit on memory usage of the executed processes in this unit. Memory usage may go above the limit if unavoidable, but the processes are heavily slowed down and memory is taken away aggressively in such cases. This is the main mechanism to control memory usage of a unit.
MemoryMax=bytes
Specify the absolute limit on memory usage of the executed processes in this unit. If memory usage cannot be contained under the limit, out-of-memory killer is invoked inside the unit. It is recommended to use MemoryHigh= as the main control mechanism and use MemoryMax= as the last line of defense.
I use a similar script to prevent my torrent client from pushing more useful data out of the page cache.You can put resource limits on basically anything (see systemd.resource-control(5)).
In Linux I was never been able to generate a similar behaviour. No matter how aggressive I set the swappiness, the system didn't began to swap until it was practically locked up with the mouse pointer lagging.
Windows reserves some memory for its functions, for example the UI (maybe the fact that the GUI is in the kernel helps), and thus doesn't have these kind of problem, especially with an SSD (and I know that page file on SSDs is not ideal... but had a disk for 5 years swapping on it and never had a problem).
The mouse cursor did move smoothly throughout, though, as I recall.