Non-uniform memory access meets the OOM killer
rachelbythebay.com
rachelbythebay.com
Also, while I'm talking prophylactics, if you have (which you should) monitoring and alerting in your production environment, it seems like there should be an alert for whenever the OOM killer activates. Assuming you are allocating resources carefully enough that you expect everything to fit, if it fires, it's almost always a sign that things are not going according to plan and need to be investigated sooner rather than later.
I can't parse this; could you clarify? What exactly do you mean by "relinquish compute access right before anyone would check up on it"?
The idea is to give IO-bound tasks a little more priority in comparison. IO-bound tasks generally run faster if they can get another IO operation started ASAP after the last one completes, and since they're not going to hog the CPU anyway, giving them relatively higher priority is one way to do this.
Anyway, the point is that the trick works by gaming the system of priority levels. If you're at the tail of one priority level, you may be off than at the head of another.
This is 100% a feature. If you care at all about memory access latency, you want to remain local to the NUMA node. Foreign memory access is significantly slower. If you have NUMA enabled and your applications are not NUMA aware, and there are shared pages being access by applications running on both nodes, the NUMA rebalancing can actually cause even worse performance as it constantly moves the pages from one node to the other.
Any application that cares about memory access latency should 100% be written to be NUMA aware, and if it is not, you should be using numactl to bind the application to the proper node.
This also goes for PCI-E devices (including nvme drives!) as they are going to be bound to a NUMA node as well. If you have an application that is accessing an nvme volume, or using a GPU, you should 100% make sure that it is running on the same node as the pci-e bus for that device.
Most environments I've worked with have to define an instance size (in memory and CPU), and determine how many parallel threads/processes will run on it. Plus you need to determine when and how to scale up to more instances. To reduce costs, the goal is to 100% utilization, but also with the capability deal with spikes in traffic an workload, and all with an acceptable error rate.
Unfortunately, doing this type sizing/scaling analysis is incredibly difficult. The opaque effects of the OOM make it even more difficult. I'm sure the OOM uses a deterministic algorithm, but it's complex enough that most don't know it, or handle for it. In a server environment, if the OOM kills a service, your app and all other services are likely hosed. It would be far more preferable if the OOM had a straightforward, consistent, and deterministic method to dealing with low memory. This way programmers would know to look out for it, and could handle it more consistently.
You can do that by tuning Linux to have this behavior if you like it. In a Unix system, I don't thing this is the most desirable behavior in the general case...
vm.overcommit_memory=2
And then the kernel will no longer do overcommit.
For systems where deterministic behavior is more important than flexibility (for me: all of them, for others: their choice) you're better off disabling overcommit.
Sadly there is no good (& simple) solution under most Unixes (that's a problem that only affects long running and big processes, but those are often the main applications of some embedded systems). I'm not really found of vfork now that we use threads more often (well, at least on Linux it seems other threads are not suspended, but it seems that this is not a Posix guarantee)... posix_spawn could be a solution if implemented with some dedicated kernel support (or always implemented through well-behaved existing ones).
Should this be considered poor behavior or design?
It's been quite a while since I wrote explicit fork/ exec code, but wouldn't a better approach be to have a small master process that spawns off the necessary children and then either links them up or mediates communication?
I mean, on a unix-like system, init is ultimately the spawner of everything else and it's not a particularly large process.
As for init being small, I leave you the responsibility of your words, now that we live in a systemd world :p
In that case OOM killer is a pretty nice solution compared to the alternatives.
The scenario I described above is HPC clusters in a university environment. The problem is students running programs that are poorly written. I'd rather reboot the node and tell them to fix their code than deal with trying to accommodate their careless / naive programming.
It turned out that the OOM killer was triggering because I was filling up memory with compiler invocations.
I was really proud of that bug.
No let's do overcommit (malloc always works) and OOM-kill some random process when under memory pressure!
void* xmalloc(size_t s) {
void *p;
if (!(p = malloc(s)) {
panic(“OOM”);
}
return p;
}
If your process ran out of RAM, you get to quit. Why offload it on some other random process? This is how your database process runs out of memory, and your web workers get killed (or vice versa). In either case your system isn’t usable, but one is harder to debug.In fact having the above in stdlib would have been just as effective at fixing thousands of little bugs and would have collectively saved us all thousands if not more man hours of developing the OOM killer and then raging online about it.
This can happen much later in time.
Memory allocated by malloc is not necessarily there by the time you need it.
Another issue is that a fork will simply copy the task state and set all the pages to 'copy on write' in the clone. That way the fork can execute very rapidly and only when the child process starts writing pages is there some actual memory overhead.
On my current laptop, pulseaudio has 2.7 GB virtual memory and 17 MB physical memory. If I open a new Firefox tab, I would much prefer for pulseaudio to die than for Firefox to die.
(You might say, "Well, pulseaudio shouldn't allocate 2.7 GB virtual memory," to which I'd say, "Sure, if you'd rather spend the thousands of hours fixing pulseaudio and every other program like it, go for it.")
Also, "panic" isn't a standard C function - are you envisioning a panic that unwinds or one that doesn't? If it unwinds, what happens if you need to allocate memory during the unwind to save state? If it doesn't, why is it a better state of the world for Firefox to crash periodically and require me to recover my work, without even giving it the option to synchronize its state?
heh, anything compiled with GHC allocates a TERABYTE of virtual memory, good luck running that w/o overcommit
This is how iOS has handled things since version 2.0: https://developer.apple.com/documentation/uikit/uiapplicatio... > It is strongly recommended that you implement this method. If your app does not release enough memory during low-memory conditions, the system may terminate it outright.
(Signals for asynchronous conditions are an awkward interface because they can interrupt you between any two assembly instructions. You're not able to release memory in the handler itself; you have to set a flag that gets handled by the main loop. So eventfd makes sense here. I'm assuming iOS is doing something similar by queueing an Objective-C method call. Signals make a lot more sense for segfaults and the like, where you're being interrupted at the exact instruction that isn't working and you need to handle it before executing any more instructions.)
According to [1], the OOM killer is a consequence of the fork syscall, and can't be removed without breaking backward compatibility.
[1] https://drewdevault.com/2018/01/02/The-case-against-fork.htm...
Yes, there are alternative APIs that avoid the need to temporarily hold on to that memory in both processes, but the idiom is still incredibly common. And it’s not unusual for server-style processing to fork without an exec and largely share their parent’s allocation via copy-on-write. That defers allocation until there’s a page fault on a simple memory write, which doesn’t have a particularly helpful mapping to C if you want to return failure to the copying process.
If I have overcommit off, and a program that allocates more than half of my physical memory (RAM + swap), it can't fork - even if it's going to immediately exec a tiny program. If I have overcommit on, it can fork and exec, and nobody gets OOM killed.
The tradeoff is that, if I have overcommit on and the child process starts modifying every page instead of exec'ing, then the OOM killer triggers. The bet is that a) probably the child process will be killed, since it has the most physical memory in use (assuming the parent process has not touched every page it has allocated), so this isn't worse than preventing the child process in the first place, and b) this situation is rare.
The basic idea isn’t that having malloc return NULL on failure is worse than killing a random process when any process runs out of memory. The idea is that programs often ask for memory that they don’t need, or don’t need right away.
When a program actually uses too much memory, there isn’t a way to signal failure, so the OOM gets involved. If this is wrong, turn off overcommit. Unfortunately, overcommit is a system-wide setting, so individual programs can’t opt in or out.
I recently saw an interesting Windows VM feature, where a process can indicate that some pages can be sacrificed to the OS if needed, and another function to try to get them back unmodified if by chance they have not been. I don't know if similar Linux syscalls exists, but I found the principle interesting.
You can cancel the free request by performing a write, which is also a bit of an awkward API.
Consider a recursive function. The amount of memory it could need is essentially infinite, though obviously stuff breaks when the address space wraps around or when the stack collides with a different allocation.
So, do we just prohibit recursion? Do we need to have a privileged compiler (making signed executables) that refuses to generate code unless the stack usage can be proven at compile time?
Pretty much, if recursion is supported, you have overcommit on some level. For example, if we force developers to specify a stack size, we still have a form of overcommit. We've just moved it to the developer making promises that he probably can't verify to be safe.
This is what MISRA C does
> Do we need to have a privileged compiler (making signed executables) that refuses to generate code unless the stack usage can be proven at compile time?
This is roughly what the Linux kernel does (I believe it's more a set of checks that aren't formally proven)
So yes, if you want to have run a process reliably, you shouldn't be doing any potentially unbounded recursion.
> We've just moved it to the developer making promises that he probably can't verify to be safe.
I don't agree, programs get a pretty large amount of stack space by default, and as long as you aren't doing any crazy amount of alloca or recursion (which is in fact the case for most programs), you'll be fine. They may not be formally verified to be safe, but practically the are. OTOH, the OOM killer can take you down regardless of what you're doing, based on what other processes are doing, so developers can have no assurance whatsoever that their programs won't suddenly die. Anecdotally, I've only even seen stack overflows due to infinite recursion bugs that are caught during development, but I see OOM-related problems not at all infrequently.
It still does result in killed processes by default - but you can set up a segfault handler with an alternate (preallocated and non-overcommitted) stack which can handle it gracefully, unlike the OOM killer's SIGKILL. Or you can write your code so that each function call checks that stack space is available before using it. Go and Rust both used to do this, via a function called __morestack, because they both used to support segmented stacks. GCC and LLVM both support this scheme. (And for a while Rust still called __morestack for the purpose of printing a nicer error message without having to set a segfault handler: https://ldpreload.com/blog/stack-smashes-you )
But "overcommit" isn't a synonym for "something that kills processes when they exceed their memory allocation."
If you run an ordinary Linux program, exactly how much memory should be allocated to the stack? That question is essentially unanswered. At the ABI level, considering just the ELF binary itself, there is no limit in place. (no fixed size) This is a promise to provide unlimited memory.
It is obvious that giving unlimited memory to every running process is not possible.
If we were to eliminate overcommit and attempt to boot the system, the init process could not legitimately start. That process implicitly requests an unlimited amount of memory for the stack, which is obviously unavailable, so it can not be started.
Not at all:
(initramfs) grep stack /proc/1/limits
Max stack size 8388608 unlimited bytes
(initramfs) grep stack /proc/1/maps
7ffca6fb3000-7ffca6fd4000 rw-p 00000000 00:00 0 [stack]
init can request more than 8 MB if it needs it, but the binary starts with 128 kB allocated (and more faultable if you hit the guard page) and an 8 MB limit. It can change its own rlimits, sure, but it needs to do that within its existing 8 MB stack.It would be perfectly straightforward to permit setrlimit to fail if you request too much of a stack, and disable overcommit entirely. So init can raise its stack by some amount, but not more than the available physical memory on the machine.
See also public header file <linux/resource.h>:
/*
* Limit the stack by to some sane default: root can always
* increase this limit if needed.. 8MB seems reasonable.
*/
#define _STK_LIM (8*1024*1024)
> At the ABI level, considering just the ELF binary itself, there is no limit in place. (no fixed size) This is a promise to provide unlimited memory.The ABI also provides no limit on the size of bss, or on the size of the program itself. This is hardly a promise to provide unlimited memory for global variables or program text or anything else. The amount of memory is finite and unspecified by the SysV ABI and left up to implementations.
It's too much because even 8 MiB is excessive without overcommit.
It's not enough because nothing about the ELF binary even bothers to claim that any particular amount of stack is enough.
I really wouldn't call "8 MiB" an answer to how much stack space should be allocated or committed.
As of today, you can get a two NUMA nodes processor (AMD threadripper 1900X) for as little as $449.