Linux Memory Overcommit (2007)
opsmonkey.blogspot.com
opsmonkey.blogspot.com
Kubernetes does this through node pressure eviction but it is pretty easy to hook into the pressure stall information and have the application handle this as well (for example, start returning 429 HTTP responses when PSI goes over a certain level).
At the end of the day the optimal overcommit is workload dependent — good metrics and an iterative approach is needed.
That's a basic fact of life for CPU cycles even for consumers on their machines, we call that just CPU priorities or so. But it's also useful for the other resources; but there it's a bit more disruptive, because we generally need to disrupt a program to eg claim back the disk space it uses.
vm.overcommit_memory == 0 heuristic overcommit
vm.overcommit_memory == 1 (full overcommit) allows allocating more memory than there is ram + swap.
vm.overcommit_memory == 2 never overcommit. There must be enough physical ram or virtual swap to allocate the memory.
Still good for flushing out scenarios where i.e. your build process needs 64GB of memory for approximately 500ms and then drops back down to reasonable levels, due to parallelism
By killing your process and/or by failing fork()s.
In general, malloc() in Linux will never fail for small allocations.
That's in kernel-space, not user-space processes. If you have overcommit disabled, it will be user-space processes making syscalls mmap() or sbrk() to get more heap memory (in which small allocations reside), that fails with an error return code, I assume. But I guess it could also fail to auto-grow the main thread stack, and presumably get a segfault. If the kernel needs to allocate more memory for a fork(), that syscall can fail. But if the kernel needs to allocate memory for some other internal purpose, probably back to the OOM killer, I guess ... hopefully this would be very rare?
Nope. No overcommit simply means that the OOM killer will be invoked earlier. The errors likely won't happen on the synchronous path due to the way glibc() maps the pages.
Long ago, I tried to write an OOM-resilient container library, and I wanted to test it. It worked as expected on Windows, but on Linux I was not able to get NULLs from malloc(). No amount of MLOCKALL and other tricks helped.
glibc allocator, like pretty much any other allocator, maps pages in advance without touching them. It's quite likely that you'll get OOM killed when you first try to touch them, rather than during mmap() for new pages.
(This is documented in https://www.kernel.org/doc/html/latest/mm/overcommit-account... )
Doesn't vm.overcommit_memory == 2 disable overcommit? That's the situation we're talking about...
Or are you suggesting that there is some other failure path that results in my process getting killed if my own malloc() fails to get pages? (That seems wrong. Should simply return NULL in that case, no?)
It can be your process or some other process, depending on the OOM killer weights and heuristics.
> Or are you suggesting that there is some other failure path that results in my process getting killed if my own malloc() fails to get pages? (That seems wrong. Should simply return NULL in that case, no?)
No, malloc() in Linux will never return NULL for small allocations. What happens is that glibc will try to get more pages for its heap by calling mmap(). The call to mmap() will result in the OOM killer waking up and freeing some memory.
It then can kill some other process, and your process mmap() will succeed, or it can kill your process, and in this case mmap() never returns.
The malloc(), calloc(), realloc(), and reallocarray() functions
return a pointer to the allocated memory, which is suitably
aligned for any type that fits into the requested size or less.
On error, these functions return NULL and set errno.
Problem is that thanks to overcommit malloc sort of always succeeds, and then many (most?) code doesn't check that malloc successfully returned and proceeds to march on towards a later obscure failure, hence below:> It's not an orderly 'abort: malloc failed' or process termination like I would've expected. Instead, you get weird intermittent failures that don't look like OOMs
in my experience it is exactly the opposite. unless you swap onto ssds, swapping will grind the whole system to a halt, while without swap the oom-reaper will kill the culprit and you can immediately work on the system to check/fix things. result also depends on your applications of course.
For a process you can use setrlimit() (ulimit command)
Implicit overcommit is really only necessary for fork and general backwards compatibility.
the more you read about linux oom killer, overcommit, linux memory accounting heuristics, copy-on-write, fork+exec as a mechanism to spawn an unrelated process, the more surprising it is that anything works at all even some of the time.
However this is very rare to happen in practice -- stack growth isn't very frequent in typical applications; so usually a malloc() will fail first in low-memory situations.
A Windows program can avoid the risk of getting terminated on stack growth by specifying equal commit and reserve sizes for the stack. So in least in theory, it's possible to write a windows program that is reliable in low memory situations.
The whole crux of overcommit is that the affected process doesn’t get a chance to gracefully fall back or shutdown on OOM conditions at defined places in the code (like at invocations of malloc).
"We run a pretty java-heavy environment, with multiple large JVMs configured per host. The problem is that the heap sizes have been getting larger, and we were running in an overcommitted situation and did not realize it. The JVMs would all start up and malloc() their large heaps, and then at some later time once enough of the heaps were actually used, the OOM killer would kick in and more or less randomly off one of our JVMs."
I know that 2007 may have been different times, but I'd argue that max heaps for all your JVMs running on a system probably shouldn't exceed around 88% of the total system memory. (percentage goes up as the total system memory goes up from 128GB -> 256GB -> 512GB)
An alternative interpretation of the issue would be that if your OS allows a single process to overcommit more than the sum of total RAM+SWAP on the system, then maybe it is it who sucks.
Not at all. Linux supports the clone system call which allows applications complete control over this. Linux user space chooses not to take advantage of it.
For me, for a modern large CPU implementation, it simplifies a lot the userland software stack to be able to book huge virtual address spaces in userland, which are populated on page faults. If really needed, I can "release" some memory ranges of this memory virtual address space (linux munmap has some options related to just that).
I do realize I am working more and more in a 'infinite memory model' (modulo some "released" ranges from time to time).
For instance, I would be more that happy to book a few 64GiB virtual address ranges on my small 8GiB workstation.
I am talking about anonymous memory, not file backed mmaping which is another story.
Imo a good compromise to this day is to overcommit RAM with swap as fallback for correctness even if performance falls of a cliff if you use it. All of that is on top of resource limits and usage monitoring.
One has to doubt...
Which clearly isnt true. If OS development worked like userspace/web dev we'd all be in much more pain.