kicking the can down the road means there isn't any longer a reasonable correction (failing the allocation), but instead we get to drive around randomly trying to find something to kill.
this is particularly annoying if you are running a service. there is no hope for it to recover - for example by flushing a cache. instead the OS looks around - sees this fat process just sitting there, and .. good news, we have plenty of memory now.
That's what Linux does. But how do you return ENOMEM when the copy does happen and now the system is out of memory? Memory writes don't return error codes. The best you could do is send a signal, which is exactly what the OOM killer does.
You as swap not because you need to actually use it, but on order to be able to guarantee there is enough memory available if the worst case scenario happens. In normal circumstances, swap should never really be utilised.
Swap is actually there so anonymous pages can be evicted, and is often used long before there is memory contention, during normal operation.
Not having swap means that only file-backed pages can be evicted. During memory contention, this can cause thrashing. During "normal" operation, it degrades performance.
Swap only being used when memory runs out, as kind of "emergency RAM", is a very widespread misunderstanding.
I don't agree. There're plenty of dormant virtual memory pages which will never be used. Keeping them in RAM is wasting precious resources.
(I note that Windows has a different approach, with "reserve" vs "commit", but nobody regards that as a preferential reason for using Windows as a server OS)
Don't allow it. Fork the process with read only pages except for the ranges passed to fork(). Count read-write pages as used memory by the child process.
If the forked process wants to write to a page that's read-only it'll have to do a system call to turn it read-write. That call can then fail if there's not enough free memory to copy the pages.
Of course, kinda hard to fix this now...
This doesn't work because of the horrible fork/exec design. If I am a huge process and I want to run `ls`, I will first have to clone myself using fork(), which may trigger an OOM, even if the first action of my clone would have been exec(), ignoring all of that memory.
I am always surprised that no one has added a sane 'spawn process' primitive to replace fork/exec. Especially since fork() without exec() only really works in single-threaded processes.
Probably you'd get people just marking as read-write and returning in sigsegv handlers, which I'm sure has great security properties... OTOH, at least there's an opportunity to deny the remap in the handler and get a decent crashdump from the program or a sliver of hope for managing the situation.
Well, yes, if you can break backwards compatibility you can do anything. Except run all the existing software.
Overcommit was godsent in the times of expensive memory and when people used virtual memory on disk (so it will spill low use memory pages there instead of the kill). Of course these days with abundance of cheap memory and people not configuring virtual memory any more we get the situation you describe.
> If you actually hit swap the system effectively deadlocks anyway
that depends. In many cases in the past the options would be either with swap and thus slow or pretty much not at all. And again these days there is so much memory that there is always a way to avoid the swapping. Though i've met funny situations in recent years like when a several terabyte sized database process would get killed by the OOMKiller on a machine with overcommit left on and no swap configured.
Both styles of memory allocation have their uses, and their drawbacks, but please understand them before declaring many OS designers as stupid and dumb.
I know it can seem like a distinction without a difference, but I think it's fair to critique work. I think they could have been more articulate and considerate towards the designers, and have some empathy that many people have worked very hard on it, and did their best in the problem/solution space they were working with.
I just think it's important that people stay objective on if the quality of the people themselves is in question. It's not ok if it is, but I don't think it was here.
now people that don't check how malloc() returns being blanketly labled as lazy is an example of it being about people. people don't ignore the possibility of a nullptr return from malloc because they're lazy. They ignore it because it's hard -- even if you do catch it, there's very little that you can actually do. You can't dynamically allocate...so I hope you've got enough stack space to do what you've gotta do. And even then, you have to be able to propagate that there's no memory all the way up the stack as it unwinds...and check every single allocation.
The cpp world is mildly better in that it can throw std::bad_alloc...but if you have something that winds up doing an allocation in a destructor, I imagine that's not a fun time.
Most of the time, there's not really a better thing to do than crash -- and there's not a lot of incentive to put any work into it. It's not something that should be happening on any sort of regular basis.
Also, the problem with Linux is not having overcommit, but notv being able to choose when to overcommit and when not. Windows makes that easier, AFAIU.
Vfork is usually a better solution.
Edit: and yes, I would love to be able to disable overcommit per process.
I think the biggest downside to vfork is (from the Linux manpage) "the behavior is undefined if the process created by vfork() [...] calls any other function before successfully calling _exit(2) or one of the exec(3) family of functions."
So strictly speaking I think doing any of those operations documented by posix_spawn yourself between vfork and execv is undefined behavior. In practice, I believe it's fine, and on glibc posix_spawn is apparently written in terms of vfork, but libc is allowed to make assumptions that portable programs shouldn't.
Story time: I once worked on a Unix-based system with no MMU. The fork implementation did something insane: it looked for things that were possibly pointers (four-byte aligned memory locations which can be interpreted as valid physical memory locations on this system allocated to that program) and adjusted them. It was kind of like a conservative GC in that it treated anything that looked like a pointer as if it were a pointer. But no reasonable person would call modifying memory that may or may not be pointers to be "conservative". Amazingly, it worked most of the time, but sometimes strings got corrupted, and as a workaround there were a bunch of places where string buffers were fully zeroed where just a NUL byte would otherwise do.
This was early in my career, and I didn't design the system anyway. If I were working on it today I'd remove the fork implementation and make everything use vfork and/or posix_spawn instead. Apparently that's what posix_spawn was made for.
Another weird problem with the lack of MMU: the compiler also didn't use register-based addressing (what's the term? like position-independent code but for the data segment?), so global variables were truly global, not just global to the process. A bunch of code needed to be "deglobalized" to deal with this. I wanted to improve the compiler, but long story short I got offered another job first.
Interesting, thanks. That didn't come up on our system—not only did we not use shared libraries but we also linked the whole system into one binary, kernel and all. Link times were atrocious in combination with identical code folding and big VLIW sentences, but it worked.
In other words, you write your C code inside a C string instead of in C.
the flexibility of a separate fork() then exec() means you can set up the initial state of a new process exactly as you want, by doing whatever work is needed between the two calls. If you merge them into one, then you will never be able to encapsulate all of that.
“After a fork() in a multithreaded program, the child can safely call only async-signal-safe functions (see signal-safety(7)) until such time as it calls execve(2)“
For example, you can't do any dynamic memory allocation or call a function that may allocate.
Standard description
(From POSIX.1) The vfork() function has the same effect as fork(2), except that the behavior is undefined if the process created by vfork() either modifies any data other than a variable of type pid_t used to store the return value from vfork(), or returns from the function in which vfork() was called, or calls any other function before successfully calling _exit(2) or one of the exec(3) family of functions.
The proper semantics for starting a new process is something like posix_spawn or Win32 CreateProcess, i.e., you specify an executable image to start.
:) Ah.. I guess when "I was younger" (TM) I was also so dogmatic on most things computer-related.
Anyways, when we need to do something a bit more complex than equivalent of system(), then it quickly becomes evident, that in many cases we need to prepare ourselves for the future execve().
Here's the list of syscalls/c-funcs, which are called in two random projects I maintain, after fork() and before execve() (or execveat() or fexecve()).
alarm(0); /* disable alarms */
setenv(); /* A couple of required envs, like MALLOC_PERTURB_ or MALLOC_PERTURB_ */
prctl(PR_SET_DUMPABLE, 1) /* regarding ptrace()-attach */
syscall(__NR_personality, ADDR_NO_RANDOMIZE); /* disable ASLR for debugging, if needed */
socketpair() /* reliable execve success detection, some form of witchcraft */
setpriority()
prctl(PR_SET_PDEATHSIG, SIGKILL); /* die upon parents death */
setrlimit(); /* set of reset rlimits */
lseek(fd, 0, SEEK_SET); /* rewind input file for this specific subprocess */
/* prepare arguments (argv) for execve dynamically */
sysconf(_SC_NPROCESSORS_ONLN); pthread_setaffinity_np(); /* pin subprocess to a list of CPUs */
/*
LOTS of functions here
if we wanted to use net/process/mount namespacing
e.g:
assigning IP adddresses to interfaces
creating custom views of the filesystem tree
modifying capability sets
*/
open("/proc/self/oom_score_adj"), write(), close(); /* adjustment of oom score */
open("/proc/self/fd", O_DIRECTORY); getdents(); fcntl(F_GETFD); fcntl(F_SETFD, FD_CLOEXEC); close() /* closing fds upon exec */
setsid(); /* new session */
sigprocmask(empty_set); /* reset signal mask */
open("/dev/null"); dup2(null, 0..1); /* close fd 0,1,2 */
prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0)
prctl(PR_SET_SECCOMP, SECCOMP_MODE_FILTER); /* application of sandboxing */
and finally execv() or execveat()
Granted, the projects are maintain are probably more on the heavy side of things, when it comes to process manipulation, before execve, but putting all of that in some control structure, would be down to impossible for me. Such structure would have to be so extensible, that it'd have to be some form of VM I guess effectively. So.. having ability to simply call a couple of syscalls from the context of a regular new process, and before execve() is quite good here.Sure.. maybe we should have some simple form of fork/execv, for those who want to call system() or popen() and not hit the memory overcommit related crashes.
But not as a replacement, rather a new syscall. Even so, debugging while a process creation/execution failed would be madness, given that you simply would get EINVAL, and the failure could be related to any of dozen parameters in a process creation control structure.
int process_handle = spawn("/bin/true", SP_PAUSED); // create process but don't execute, and return a handle
// most API would take a process handle, thus you could do stuff (such as prctl) on the new process
unpause(process_handle);
There's a bit of a move in this direction in the Linux API with the PIDFD stuff.Also, this is all mess if fork() is called from a multi-threaded context. Esp. via clone(), which up to certain glibc() version cached getpid() result, and returned parent'd PID in the child (for performance reasons of course ;).
Main Memory => zswap (compressed memory) => swap
In this case, the pages may be logically allocated or not -- the assurance is that the data will be the value you expect it to be when it becomes resident.
Should those pages be uninitialized, the "Swapped" state is really just "Remember that this thing was all zeros."
We could do computing your way, but it'd be phenomenally more expensive. I know this because every thing we introduce to the hierarchy in practice makes computing phenomenally less expensive.
It must be viable - Windows prevents overcommit. But it has slow child-process-creation (edit: previously said "forking"), and this steers development towards native threads which is its own set of problems.
I had never previously joined the dots on the point pcwalton makes at the top of this thread. It is a dramatic trade-off.
It seems more that overcommit is a workaround for the way fork() can potentially lead to copying the entire address space into new pages, but usually doesn't. Because CreateProcess() knows more precisely how much to allocate before it returns to the child process, it can just reserve that amount and signal an error immediately if there's not enough memory+swap to back it.
(And on the other hand, Windows has a lot of legacy and backwards compatibility behavior around processes that could easily explain the slower process creation independent of the API.)
For ages cygwin had to emulate it manually by hand-copying the process.
Everything is great, up until some other process on your system does a fork bomb of an infinitely recursive program that allocates nothing on the heap. You've just got a whole lot of quickly growing stacks hoovering up your physical memory pages.
If overcommit is disabled and someone has allocated most system memory, fork() and exec() and pthread_create() etc. will theoretically fail with ENOMEM.
A bigger problem on Linux at least is that the kernel will swap out all possible memory before returning an allocation error. Even if you have not allocated any swap space, it will swap out any memory mapped files.
And even if none of your programs have explicitly mmap()ed anything, the actual application code is mapped, so it will start swapping out all code pages to disk before it refuses an allocation. And now this means that your system has become entirely unusable and will have to be hard rebooted, because it is now swapping data to and from disk on every instruction execution after every context switch. At least your CPU will get to stay nice and cool for a while.
IIUC Linux was really the first OS to make overcommit so prominent. Most systems were a lot more conservative.
I don't think disabling of overcommit implies that physical pages are mapped immediately. If caches are instantly droppable, you can use a page that's allocated but unused for cache, and drop the cache page (and zero it) when the allocated page is written to.
You'd still have all of your caches until you have memory pressure with actual data written (but of course, with overcommit, you'd drop caches then too), but if you attempt to allocate more than you have (including through fork attempts as discussed elsewhere), you get a system call failure rather than an OOM kill.
I've found the linux memory system to be far too complicated to understand for quite some time, compared to what's documented in, for example, The Design and Implementation of The FreeBSD Operating System, for a more comprehensible system.
(I’m talking in general here - not Linux specifically)
Everything I'm describing is about a busy server with heterogenous workloads of specific types.
Couldn't the system just reserve the pages for future use by the application but still use them for caching until the application actually tries to use them?
Somehow Solaris manages just fine.
And don't forget that swap memory exists. Ironically, using overcommit without swap is asking for trouble on Linux. Overcommit or no overcommit, the Linux VM and page buffer systems are designed with the expectation of swap.
fork and malloc can also fail in Linux even with overcommit enabled (rlimits, but also OOM killer racing with I/O page dirtying triggering best-effort timeout), so Linux buys you a little convenience at the cost of making it impossibly difficult to actually guarantee behavior when it matters most.
[1] https://docs.oracle.com/cd/E26505_01/html/816-5167/vfork-2.h...
Some other type of process like an interpreter that can subshell out doesn't know how big the allocation is going to get, would have to pre-fork early on.
In this way, you wouldn't "need" overcommit and the Linux horror of OOM. Well, perhaps you don't need it so badly. Programs that use sparse arrays without mmap() probably need overcommit or lots of swap.
https://www.kernel.org/doc/Documentation/vm/overcommit-accou...
We didn't actually want to fork anything and share gigabytes of virtual memory with the child process, we wanted to spawn an almost entirely independent process to do something and report results, but that got implemented under the hood by fork.
Spawning processes is one area where Windows is more elegant than linux: windows offers spawn. Apparently macos and solaris implement a posix_spawn that avoid the complications of fork/exec.
linux offers posix_spawn, apparently which may may or may not call fork under the hood depending on which libc you're using. If libc implements posix_spawn by calling fork then you're back in the same mess with linux heuristic memory accounting and overcommit. E.g. old versions of glibc will fork when you posix_spawn, newer versions of glibc may vfork . musl apparently will always vfork.
It looks like cpython's subprocess.Popen was patched in python 3.8 to detect some cases where posix_spawn can be used -- it reads as if it will only kick in on linux if it detects a sufficiently new version of glibc: https://github.com/python/cpython/blob/main/Lib/subprocess.p...
edit: Python 3.10 now supports using vfork for linux inside subprocess: https://bugs.python.org/issue35823
docker run --rm -it --entrypoint=/bin/sh python:3.9-alpine
# apk add strace
# strace python -c "import subprocess; subprocess.run(['ls', '-l'])" 2>&1 >/dev/null | grep fork
fork() = 88
docker run --rm -it --entrypoint=/bin/sh python:3.10-alpine
# apk add strace
# strace python -c "import subprocess; subprocess.run(['ls', '-l'])" 2>&1 >/dev/null | grep fork
vfork() = 15
edit 2: here's a similar tale from go, replacing use of fork in fork/exec:https://github.com/golang/go/issues/5838
https://go-review.googlesource.com/c/go/+/37439/
https://about.gitlab.com/blog/2018/01/23/how-a-fix-in-go-19-...
The manual page is quite informative: man clone
years ago when i first hit this in production, we ended up working around it by rewriting our application code to use a pure-python library that did the equivalent thing as the separate command line tool we were trying to launch.