Both styles of memory allocation have their uses, and their drawbacks, but please understand them before declaring many OS designers as stupid and dumb.
Also, the problem with Linux is not having overcommit, but notv being able to choose when to overcommit and when not. Windows makes that easier, AFAIU.
Vfork is usually a better solution.
Edit: and yes, I would love to be able to disable overcommit per process.
I think the biggest downside to vfork is (from the Linux manpage) "the behavior is undefined if the process created by vfork() [...] calls any other function before successfully calling _exit(2) or one of the exec(3) family of functions."
So strictly speaking I think doing any of those operations documented by posix_spawn yourself between vfork and execv is undefined behavior. In practice, I believe it's fine, and on glibc posix_spawn is apparently written in terms of vfork, but libc is allowed to make assumptions that portable programs shouldn't.
Story time: I once worked on a Unix-based system with no MMU. The fork implementation did something insane: it looked for things that were possibly pointers (four-byte aligned memory locations which can be interpreted as valid physical memory locations on this system allocated to that program) and adjusted them. It was kind of like a conservative GC in that it treated anything that looked like a pointer as if it were a pointer. But no reasonable person would call modifying memory that may or may not be pointers to be "conservative". Amazingly, it worked most of the time, but sometimes strings got corrupted, and as a workaround there were a bunch of places where string buffers were fully zeroed where just a NUL byte would otherwise do.
This was early in my career, and I didn't design the system anyway. If I were working on it today I'd remove the fork implementation and make everything use vfork and/or posix_spawn instead. Apparently that's what posix_spawn was made for.
Another weird problem with the lack of MMU: the compiler also didn't use register-based addressing (what's the term? like position-independent code but for the data segment?), so global variables were truly global, not just global to the process. A bunch of code needed to be "deglobalized" to deal with this. I wanted to improve the compiler, but long story short I got offered another job first.
Interesting, thanks. That didn't come up on our system—not only did we not use shared libraries but we also linked the whole system into one binary, kernel and all. Link times were atrocious in combination with identical code folding and big VLIW sentences, but it worked.
In other words, you write your C code inside a C string instead of in C.
the flexibility of a separate fork() then exec() means you can set up the initial state of a new process exactly as you want, by doing whatever work is needed between the two calls. If you merge them into one, then you will never be able to encapsulate all of that.
“After a fork() in a multithreaded program, the child can safely call only async-signal-safe functions (see signal-safety(7)) until such time as it calls execve(2)“
For example, you can't do any dynamic memory allocation or call a function that may allocate.
Standard description
(From POSIX.1) The vfork() function has the same effect as fork(2), except that the behavior is undefined if the process created by vfork() either modifies any data other than a variable of type pid_t used to store the return value from vfork(), or returns from the function in which vfork() was called, or calls any other function before successfully calling _exit(2) or one of the exec(3) family of functions.
I know it can seem like a distinction without a difference, but I think it's fair to critique work. I think they could have been more articulate and considerate towards the designers, and have some empathy that many people have worked very hard on it, and did their best in the problem/solution space they were working with.
I just think it's important that people stay objective on if the quality of the people themselves is in question. It's not ok if it is, but I don't think it was here.
now people that don't check how malloc() returns being blanketly labled as lazy is an example of it being about people. people don't ignore the possibility of a nullptr return from malloc because they're lazy. They ignore it because it's hard -- even if you do catch it, there's very little that you can actually do. You can't dynamically allocate...so I hope you've got enough stack space to do what you've gotta do. And even then, you have to be able to propagate that there's no memory all the way up the stack as it unwinds...and check every single allocation.
The cpp world is mildly better in that it can throw std::bad_alloc...but if you have something that winds up doing an allocation in a destructor, I imagine that's not a fun time.
Most of the time, there's not really a better thing to do than crash -- and there's not a lot of incentive to put any work into it. It's not something that should be happening on any sort of regular basis.
The proper semantics for starting a new process is something like posix_spawn or Win32 CreateProcess, i.e., you specify an executable image to start.
:) Ah.. I guess when "I was younger" (TM) I was also so dogmatic on most things computer-related.
Anyways, when we need to do something a bit more complex than equivalent of system(), then it quickly becomes evident, that in many cases we need to prepare ourselves for the future execve().
Here's the list of syscalls/c-funcs, which are called in two random projects I maintain, after fork() and before execve() (or execveat() or fexecve()).
alarm(0); /* disable alarms */
setenv(); /* A couple of required envs, like MALLOC_PERTURB_ or MALLOC_PERTURB_ */
prctl(PR_SET_DUMPABLE, 1) /* regarding ptrace()-attach */
syscall(__NR_personality, ADDR_NO_RANDOMIZE); /* disable ASLR for debugging, if needed */
socketpair() /* reliable execve success detection, some form of witchcraft */
setpriority()
prctl(PR_SET_PDEATHSIG, SIGKILL); /* die upon parents death */
setrlimit(); /* set of reset rlimits */
lseek(fd, 0, SEEK_SET); /* rewind input file for this specific subprocess */
/* prepare arguments (argv) for execve dynamically */
sysconf(_SC_NPROCESSORS_ONLN); pthread_setaffinity_np(); /* pin subprocess to a list of CPUs */
/*
LOTS of functions here
if we wanted to use net/process/mount namespacing
e.g:
assigning IP adddresses to interfaces
creating custom views of the filesystem tree
modifying capability sets
*/
open("/proc/self/oom_score_adj"), write(), close(); /* adjustment of oom score */
open("/proc/self/fd", O_DIRECTORY); getdents(); fcntl(F_GETFD); fcntl(F_SETFD, FD_CLOEXEC); close() /* closing fds upon exec */
setsid(); /* new session */
sigprocmask(empty_set); /* reset signal mask */
open("/dev/null"); dup2(null, 0..1); /* close fd 0,1,2 */
prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0)
prctl(PR_SET_SECCOMP, SECCOMP_MODE_FILTER); /* application of sandboxing */
and finally execv() or execveat()
Granted, the projects are maintain are probably more on the heavy side of things, when it comes to process manipulation, before execve, but putting all of that in some control structure, would be down to impossible for me. Such structure would have to be so extensible, that it'd have to be some form of VM I guess effectively. So.. having ability to simply call a couple of syscalls from the context of a regular new process, and before execve() is quite good here.Sure.. maybe we should have some simple form of fork/execv, for those who want to call system() or popen() and not hit the memory overcommit related crashes.
But not as a replacement, rather a new syscall. Even so, debugging while a process creation/execution failed would be madness, given that you simply would get EINVAL, and the failure could be related to any of dozen parameters in a process creation control structure.
int process_handle = spawn("/bin/true", SP_PAUSED); // create process but don't execute, and return a handle
// most API would take a process handle, thus you could do stuff (such as prctl) on the new process
unpause(process_handle);
There's a bit of a move in this direction in the Linux API with the PIDFD stuff.Also, this is all mess if fork() is called from a multi-threaded context. Esp. via clone(), which up to certain glibc() version cached getpid() result, and returned parent'd PID in the child (for performance reasons of course ;).
Main Memory => zswap (compressed memory) => swap
In this case, the pages may be logically allocated or not -- the assurance is that the data will be the value you expect it to be when it becomes resident.
Should those pages be uninitialized, the "Swapped" state is really just "Remember that this thing was all zeros."
We could do computing your way, but it'd be phenomenally more expensive. I know this because every thing we introduce to the hierarchy in practice makes computing phenomenally less expensive.
It must be viable - Windows prevents overcommit. But it has slow child-process-creation (edit: previously said "forking"), and this steers development towards native threads which is its own set of problems.
I had never previously joined the dots on the point pcwalton makes at the top of this thread. It is a dramatic trade-off.
For ages cygwin had to emulate it manually by hand-copying the process.
It seems more that overcommit is a workaround for the way fork() can potentially lead to copying the entire address space into new pages, but usually doesn't. Because CreateProcess() knows more precisely how much to allocate before it returns to the child process, it can just reserve that amount and signal an error immediately if there's not enough memory+swap to back it.
(And on the other hand, Windows has a lot of legacy and backwards compatibility behavior around processes that could easily explain the slower process creation independent of the API.)
Everything is great, up until some other process on your system does a fork bomb of an infinitely recursive program that allocates nothing on the heap. You've just got a whole lot of quickly growing stacks hoovering up your physical memory pages.
If overcommit is disabled and someone has allocated most system memory, fork() and exec() and pthread_create() etc. will theoretically fail with ENOMEM.
A bigger problem on Linux at least is that the kernel will swap out all possible memory before returning an allocation error. Even if you have not allocated any swap space, it will swap out any memory mapped files.
And even if none of your programs have explicitly mmap()ed anything, the actual application code is mapped, so it will start swapping out all code pages to disk before it refuses an allocation. And now this means that your system has become entirely unusable and will have to be hard rebooted, because it is now swapping data to and from disk on every instruction execution after every context switch. At least your CPU will get to stay nice and cool for a while.