No let's do overcommit (malloc always works) and OOM-kill some random process when under memory pressure!
No let's do overcommit (malloc always works) and OOM-kill some random process when under memory pressure!
According to [1], the OOM killer is a consequence of the fork syscall, and can't be removed without breaking backward compatibility.
[1] https://drewdevault.com/2018/01/02/The-case-against-fork.htm...
If I have overcommit off, and a program that allocates more than half of my physical memory (RAM + swap), it can't fork - even if it's going to immediately exec a tiny program. If I have overcommit on, it can fork and exec, and nobody gets OOM killed.
The tradeoff is that, if I have overcommit on and the child process starts modifying every page instead of exec'ing, then the OOM killer triggers. The bet is that a) probably the child process will be killed, since it has the most physical memory in use (assuming the parent process has not touched every page it has allocated), so this isn't worse than preventing the child process in the first place, and b) this situation is rare.
Yes, there are alternative APIs that avoid the need to temporarily hold on to that memory in both processes, but the idiom is still incredibly common. And it’s not unusual for server-style processing to fork without an exec and largely share their parent’s allocation via copy-on-write. That defers allocation until there’s a page fault on a simple memory write, which doesn’t have a particularly helpful mapping to C if you want to return failure to the copying process.
The basic idea isn’t that having malloc return NULL on failure is worse than killing a random process when any process runs out of memory. The idea is that programs often ask for memory that they don’t need, or don’t need right away.
When a program actually uses too much memory, there isn’t a way to signal failure, so the OOM gets involved. If this is wrong, turn off overcommit. Unfortunately, overcommit is a system-wide setting, so individual programs can’t opt in or out.
I recently saw an interesting Windows VM feature, where a process can indicate that some pages can be sacrificed to the OS if needed, and another function to try to get them back unmodified if by chance they have not been. I don't know if similar Linux syscalls exists, but I found the principle interesting.
You can cancel the free request by performing a write, which is also a bit of an awkward API.
void* xmalloc(size_t s) {
void *p;
if (!(p = malloc(s)) {
panic(“OOM”);
}
return p;
}
If your process ran out of RAM, you get to quit. Why offload it on some other random process? This is how your database process runs out of memory, and your web workers get killed (or vice versa). In either case your system isn’t usable, but one is harder to debug.In fact having the above in stdlib would have been just as effective at fixing thousands of little bugs and would have collectively saved us all thousands if not more man hours of developing the OOM killer and then raging online about it.
On my current laptop, pulseaudio has 2.7 GB virtual memory and 17 MB physical memory. If I open a new Firefox tab, I would much prefer for pulseaudio to die than for Firefox to die.
(You might say, "Well, pulseaudio shouldn't allocate 2.7 GB virtual memory," to which I'd say, "Sure, if you'd rather spend the thousands of hours fixing pulseaudio and every other program like it, go for it.")
Also, "panic" isn't a standard C function - are you envisioning a panic that unwinds or one that doesn't? If it unwinds, what happens if you need to allocate memory during the unwind to save state? If it doesn't, why is it a better state of the world for Firefox to crash periodically and require me to recover my work, without even giving it the option to synchronize its state?
heh, anything compiled with GHC allocates a TERABYTE of virtual memory, good luck running that w/o overcommit
This can happen much later in time.
Memory allocated by malloc is not necessarily there by the time you need it.
Another issue is that a fork will simply copy the task state and set all the pages to 'copy on write' in the clone. That way the fork can execute very rapidly and only when the child process starts writing pages is there some actual memory overhead.
This is how iOS has handled things since version 2.0: https://developer.apple.com/documentation/uikit/uiapplicatio... > It is strongly recommended that you implement this method. If your app does not release enough memory during low-memory conditions, the system may terminate it outright.
(Signals for asynchronous conditions are an awkward interface because they can interrupt you between any two assembly instructions. You're not able to release memory in the handler itself; you have to set a flag that gets handled by the main loop. So eventfd makes sense here. I'm assuming iOS is doing something similar by queueing an Objective-C method call. Signals make a lot more sense for segfaults and the like, where you're being interrupted at the exact instruction that isn't working and you need to handle it before executing any more instructions.)
Consider a recursive function. The amount of memory it could need is essentially infinite, though obviously stuff breaks when the address space wraps around or when the stack collides with a different allocation.
So, do we just prohibit recursion? Do we need to have a privileged compiler (making signed executables) that refuses to generate code unless the stack usage can be proven at compile time?
Pretty much, if recursion is supported, you have overcommit on some level. For example, if we force developers to specify a stack size, we still have a form of overcommit. We've just moved it to the developer making promises that he probably can't verify to be safe.
It still does result in killed processes by default - but you can set up a segfault handler with an alternate (preallocated and non-overcommitted) stack which can handle it gracefully, unlike the OOM killer's SIGKILL. Or you can write your code so that each function call checks that stack space is available before using it. Go and Rust both used to do this, via a function called __morestack, because they both used to support segmented stacks. GCC and LLVM both support this scheme. (And for a while Rust still called __morestack for the purpose of printing a nicer error message without having to set a segfault handler: https://ldpreload.com/blog/stack-smashes-you )
But "overcommit" isn't a synonym for "something that kills processes when they exceed their memory allocation."
If you run an ordinary Linux program, exactly how much memory should be allocated to the stack? That question is essentially unanswered. At the ABI level, considering just the ELF binary itself, there is no limit in place. (no fixed size) This is a promise to provide unlimited memory.
It is obvious that giving unlimited memory to every running process is not possible.
If we were to eliminate overcommit and attempt to boot the system, the init process could not legitimately start. That process implicitly requests an unlimited amount of memory for the stack, which is obviously unavailable, so it can not be started.
Not at all:
(initramfs) grep stack /proc/1/limits
Max stack size 8388608 unlimited bytes
(initramfs) grep stack /proc/1/maps
7ffca6fb3000-7ffca6fd4000 rw-p 00000000 00:00 0 [stack]
init can request more than 8 MB if it needs it, but the binary starts with 128 kB allocated (and more faultable if you hit the guard page) and an 8 MB limit. It can change its own rlimits, sure, but it needs to do that within its existing 8 MB stack.It would be perfectly straightforward to permit setrlimit to fail if you request too much of a stack, and disable overcommit entirely. So init can raise its stack by some amount, but not more than the available physical memory on the machine.
See also public header file <linux/resource.h>:
/*
* Limit the stack by to some sane default: root can always
* increase this limit if needed.. 8MB seems reasonable.
*/
#define _STK_LIM (8*1024*1024)
> At the ABI level, considering just the ELF binary itself, there is no limit in place. (no fixed size) This is a promise to provide unlimited memory.The ABI also provides no limit on the size of bss, or on the size of the program itself. This is hardly a promise to provide unlimited memory for global variables or program text or anything else. The amount of memory is finite and unspecified by the SysV ABI and left up to implementations.
It's too much because even 8 MiB is excessive without overcommit.
It's not enough because nothing about the ELF binary even bothers to claim that any particular amount of stack is enough.
I really wouldn't call "8 MiB" an answer to how much stack space should be allocated or committed.
This is what MISRA C does
> Do we need to have a privileged compiler (making signed executables) that refuses to generate code unless the stack usage can be proven at compile time?
This is roughly what the Linux kernel does (I believe it's more a set of checks that aren't formally proven)
So yes, if you want to have run a process reliably, you shouldn't be doing any potentially unbounded recursion.
> We've just moved it to the developer making promises that he probably can't verify to be safe.
I don't agree, programs get a pretty large amount of stack space by default, and as long as you aren't doing any crazy amount of alloca or recursion (which is in fact the case for most programs), you'll be fine. They may not be formally verified to be safe, but practically the are. OTOH, the OOM killer can take you down regardless of what you're doing, based on what other processes are doing, so developers can have no assurance whatsoever that their programs won't suddenly die. Anecdotally, I've only even seen stack overflows due to infinite recursion bugs that are caught during development, but I see OOM-related problems not at all infrequently.