There is an OOM kill count in Linux
medium.com
medium.com
My number one recurring issue on Ubuntu is me not noticing the memory usage, running out of memory and locking up my system (forcing a hard reboot and loss of unsaved work)
EDIT: I'd preferably like an OOM handler that would freeze my system and pop up a little menu from which I could select which process to nuke
Without swap, the system lags for a couple seconds, OOM killer frees up memory and you're good to go again. The only slowdown is any pages that were kicked out from the file cache. But those quickly come back after the OOM killer does its thing.
I would view the OOm solution as a compute as cattle thinb, but here we are talking about a user desktop where the user can take the best action for themselves once they realize there's a problem.
The user doesn't decide which processes are swapped in. If the process gets CPU time and tries to access its data, that data will get swapped in.
> We are talking about a user desktop where the user can take the best action for themselves once they realize there's a problem.
You can't do that with swap, because once you realize there's a problem, you cannot even move your cursor or run commands to take any actions.
The Linux kernel OOM killer only acts if there's nothing left that can be discarded, which often happens way too late to save the system. You need a user-mode OOM-killer like earlyoom if you want to keep the system responsive.
I'm just saying turning on swap, or increasing the swap capacity does not fix the problem, and it usually makes it worse.
On my current laptop and with my current habits, I’m consuming a lot more memory, and so the technique I use to avoid OOM is simply having 40GB of RAM. (As it happens, I do actually have swap set up at present, because certain circumstances meant I wanted to hibernate it occasionally; should disable swap again now I’m back to not needing to hibernate.)
Maybe there are reasons, but it is not nearly good enough -- I frequently run out of RAM and encounter OOM kills (prefer not to deal with linux swap), usually requiring a reboot. I really wish that I could just set an upper limit on (e.g.) firefox RAM usage -- 8GB for example -- instead of its insistence on using all of the unused RAM minus a couple to several hundred MB, which does not leave enough room for memory usage spikes. There might even be a way to set this buried somewhere in the config parameters, but I could never find it. It must be technically possible to set some limit because otherwise the browser would not be able to maintain a somewhat consistent usage just below the total system RAM.
You can also use facebook's oomd, but the systemd one is basically the same and slightly better.
If you want, you can also write your own custom memory pressure handler using the recent PSI features: https://docs.kernel.org/accounting/psi.html
The other, easier, answer is to buy more ram, if you simply have more memory than you will ever use (such as having 128GiB of memory when you only regularly use 30GiB), you'll rarely run into OOM issues.
[0]: https://www.freedesktop.org/software/systemd/man/latest/syst...
Wanna emphasize this mostly because it took me way too long to think of.
Was dealing with memory pressure regularly and constantly getting frustrated… then dropped an entire $60 or something to throw another 32GB in the machine and never think about my memory again.
Like a week later when I ran out of disk space again and was struggling to make room… same realization. Spent $75 for a 1TB NVME SSD dropped in my mailbox two days later and doubled my storage.
So we’re talking $135 to double my RAM and storage and stop wasting time and brain space on this stuff.
We ran out of space, and then we realized we are a storage company.
By accident, I once wrote an infinite loop that just allocated a bunch of memory. As I ran the program my Windows system quickly became laggy and unresponsive. However, not completely. With a bit of patience I got Task Manager up and managed to kill the process and the system was back to normal within a minute.
On Linux, I experience the same as you. Everything is fine until suddenly it's completely unresponsive. Most of the time I can't even manage to shut down the system cleanly and have to hard reset.
I really think a more general solution is treating memory (and CPU time as welL!) as a scarce resource, and programs should either: (1) Deal with having severely denied resources (default behavior); (2) Use a communication protocol to negotiate memory with the system.
Negotiation could mean (a) The process freeing memory spontaneously when there's pressure (another process requires it) and it's inactive, (b) Requesting the system free memory from other processes when user requires.
This would impose more memory management burden on developers (trying to fulfill OS requests), but in turn it would make for a far better experience when there's memory contention.
Disclaimer: I have no idea how memory allocation works in detail. Maybe this is reinventing the wheel?
If you turn swap off the allocator will error out when you are out of memory. I don't think most software handles that error but all the wiring is there to do that.
I believe something like that exists on mobile platforms for memory, at least on iOS you get a message (applicationDidReceiveMemoryWarning:) when the system is memory starved and wants your app to free memory. If you don't release enough memory and the memory pressure doesn't go down, the system will start killing apps.
malloc() can fail, everyone knows this in theory, but assume you're the programmer handling the malloc() error -- you know the system is likely to have already run out of free memory -- how do you make the program fail gracefully without doing anything at all to allocate memory?
It's not even safe to call printf() since it might have to allocate a new string. There are very few things in modern code that doesn't "incidentally" allocate a couple bytes here and there.
Not to mention languages higher level than C where you don't directly call malloc() but have the language handle it for you (GC etc.). Handling OOM for those programs is even more impossible because you really have no idea whatever class/method/function you're calling is going to trigger an allocation or not.
The best you can do in those cases is try to commit critical data (if any) to disk and have the program die ASAP. Which is not that different from having the kernel kill it for you.
The OS could really just say no. But that would commit memory to a process even though it might not actually end up using it, making it unavailable for processes that need it right now. Also, memory allocation latency is higher since the OS has to really do all the bookkeeping up front. Therefore, most Linux distributions by default optimistically grant allocation requests even though there might not be enough memory available yet.
The big disadvantage is that the OS might be unable to actually grant memory. The process then gets OOM killed. This is sort of fine on servers since users usually get assigned limits. But there is something left to be desired for single-user systems.
Processes can at any time return memory to the OS when they don't need it. Usually they do that, but this has the disadvantage that they have to ask the OS again for it, which is quite slow compared to keeping it around.
Few applications are actually able to just release memory on request. It's critical user and internal application data after all. Databases and applications with garbage collectors come closest.
There are already interfaces which can do some of what you are talking about: the most common is memory mapping backed by a file, swap, or filesystem caching. This is the most common form of 'optional' memory usage, and it's generally managed by the kernel. A lot of the kind of thing you are talking about can be mapped onto this abstraction. For cases where backing by a file doesn't make sense (e.g. it's cheaper to regenerate the cached data than to read it back from disk), if you are using a mmap-like interface, you can generally mark pages as 'don't need', which means they may be freed by the kernel if there is memory pressure, but won't be wiped immediately. Again the main issue is the use-cases where this represents a significant percentage of memory usage are pretty slim.
That's actually fine on multi-user systems. A competent admin would set up tight limits and prevent one user taking over all resources, and intervenes if they mess up. On single-user machines, this limitation is not there because there is effectively only one user. And since desktop Linux is not a priority for most Linux vendors, this scenario sees comparatively little consideration.
In comparison, on Windows some applications (at least explorer.exe and the task manager) seem to have way higher priority. No matter what else is going on, the user must be able to use these application to rein in other applications. The drawback is that there is no recourse should explorer.exe ever lock up.
Windows Logon is superior to all the shells as a core Windows subsystem in that it is the one that executes the shell (usually Explorer) on boot/login and can also execute Task Manager, in addition to its primary duties of tracking and securing user logins (hence the name).
This means even if Explorer has completely froze or has crashed, you can always get to Task Manager via Windows Logon and then kill or execute another instance of Explorer (or CMD or Powershell!) or otherwise force reboot the system from Windows Logon.
Then I have stopped using swap partitions or swap files. For more than 20 years I have never used again any kind of swap, on a large variety of desktops, laptops and servers with Linux. I never had again any problem with unresponsiveness.
Nevertheless, I provision all the computers with generous amounts of DRAM, e.g. there are many years since I have last used a computer with less than 32 GB. In these conditions, OOM events are extremely rare.
While there are people who claim that there still exist cases when swap can be useful, I have never seen any evidence for this claim.
(It sometimes happens in software development and it's nice if it doesn't take your machine in swap hell as a result.)
I enabled systemd-oom, but it hasn't been necessary, because I also enabled zram as a swap device. I've seen it compress multiple gb of ram into a few hundred mb and haven't frozen my desktop since.
Make sure you use zram (or zswap) as well. May also consider enabling userspace oom handler like others already suggested, but it's less important than the rest.
Can have security implications, but the target is chosen in the typical way and screen locks should be setting a score_adj to avoid being picked.
Due to cgroupv2, systemd-oomd tends to kill entire process groups instead of a single run-away process [1]. It's perfect for servers where almost every process is part of a system service already.
For desktop usage I found this behavior unacceptable. I would sometimes get my entire X session killed when trying something dumb (which is the reason I run such a daemon in the first place), or my editor when spawning a helper subprocess...
earlyoom was the simplest daemon of the bunch I tried, and worked out of the box without special configuration.
[1] Or at least this was the state ~1yr ago.
If you can't upgrade the RAM, zram[2] is also worth a try.
As for my pragmatic solution at home: max out the physical RAM…
As far as I know, unless a browser manages its oom_score_adj, which none do, a wrapper is insufficient as a workaround and a daemon would have to do so on its behalf, because the value is inherited when forking.
[0] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
https://pallissard.net/2022/06/27/limiting_application_resou...
Edit: I like your user space oomd idea though, that would be killer
In my opinion, if your distro already integrates systemd-oom, like Fedora, it's easier to just use that. Otherwise earlyoom is the next best choice.
Also it's recommended to enable zram. It's much faster than a swap partition or file. The combination of earlyoom and zram made my system much more responsive when it's running short of memory.
(via https://news.ycombinator.com/item?id=37640389, but no comments there)