Hastening Linux process cleanup with process_mrelease
lwn.net
lwn.net
My experience of the linux OOM killer is not that its opinion differs from mine but that it has no opinion at all for a long, long time after the system is in deep trouble. The OOM killer simply does not act quickly enough to save systems. Sadly it's not customisable but 'earlyoom' (packaged for debian and probably everything else) is. I turned it on when I was messing about when debugging a badly behaved bit of software which went into a memory allocation loop and have just left it on. It's saved me a few times and I now plan to leave it on forever.
Looks like oomd is an idea along the same lines but with slightly different goals. It's not in my distro so not an easy option for me.
Fedora has earlyoom enabled by default but so far it hasn't saved me. I really need to look into configuring it. How did you get started? Man pages? Blog post?
If the results you're getting aren't expected, there's still some chance it's a bug somewhere, so you should report it against systemd, attach `journalctl -b -o short-monotonic --no-hostname` or at least ~10 minutes of logging prior to the unexpected behavior you're reporting.
-r 30 -m 5 -s 80
I run with 16gb of memory and do use swap. In practice if swap is growing at all once memory is near full, I'm in trouble and action needs to be taken.
Note that if enabled you can use SysRq keys to invoke OOM killer manually.
Given the letter they picked it seems like they're fully aware of the deficiency. As to why it was never really addressed, that's a good question. The recent le9[1] patch addresses the most annoying symptom of running out of memory by keeping some amount of RAM reserved for clean pages (like static program code) so that those won't get swapped in and out so much. That greatly improves responsiveness under low memory conditions.
There is a new patch[1] that allows to set a soft and hard minimum of RAM reserved for clean pages. This fixes the problem almost completely even under the heaviest loads.
cgroups like a silbling post mentioned can also help by setting soft limits for heavy background tasks like compile jobs. Setting a soft limit will give them as much RAM as is available, or swap them out completely when things are heavily contented, effectively pausing the processes. It requires some setting up though, so it's not a solution for all cases, but it can make sense even without the thrashing problem.
Nit: I don't think it's necessary to copy the program into anonymous memory to use mlock. You might be thinking of huge pages (transparent or otherwise), which are supported on anonymous pages and unfortunately aren't yet supported on ext4- or btrfs-backed file pages.
You'll still see that page fault thrashing but it becomes isolated to the cgroup experiencing the memory pressure. It doesn't bring down the entire system in my experience.
That is, it was my understanding that the reason a process is in an uninterruptible sleep is generally because it's waiting for a DMA to complete, and if you were permitted to interrupt it, the DMA would eventually complete and clobber who knows what. Ripping the memory away from such a process would, to my first glance, seem to encounter the same problem -- how do you stop the pending DMA (which might already be in progress, but which might also take awhile to complete.) Whatever method, it would be device dependent, which makes it impractical (who's going to retrofit all the device drivers with DMA stopping APIs, and I'm sure there are many devices that have no way to stop pending DMAs, and anyway, maybe the DMA is already in progress, only half completed.) Maybe do it at the pci level, unmap DMA buffers. But many current drivers will generally assume that they never give their devices bad bus addresses, yet now the device is attempting DMA to a bad bus address (i.e. a suddenly unmapped bus address).
Well, I've been out of the linux driver game for awhile now, so perhaps I'm missing or forgetting something. Ripping memory out from under a process with pending DMA sounds pretty sketchy to me though.
But, if dumb old me can think of this, of course the kernel developers can also, and undoubtedly did. Wonder how it really works?
It's not like it's a new issue, anyway - physical backing pages can be dropped and re-used even if the process is still running, so the reference counting always had to exist.
This isn't something userspace processes opt into. It's just how blocking read and write syscalls on filesystems work [1]. If you ever hit a bad block, you may notice that threads that touch it through the fs just hang and you can't recover or kill them. This is how it's always been on Linux and other Unix systems. I of course absolutely hate this and am 100% behind you if you're suggesting changing it.
[1] With the exception of nfs if you set the "intr" mount option.
WASM's WASI is sort of the spiritual successor, but I prefer a diversity of tactics so CloudABI and WASI should both exist.
CloudABI is the best way to save desktop computing that's not a complete Hail Mary.
How long it’s been there under my nose without me realising it? :-D
I could totally see myself running out of memory trying to open /dev/sda in emacs or doing something stupid in elisp. So killing emacs makes a lot more sense than some innocent daemon.