Process Creation in Io_uring
lwn.net
lwn.net
I wonder if the best solution lies somewhere in the vicinity of "fork but only copy a small part of the address space" -- rather than copying the entire address space as in fork (only to use a tiny portion and throw away the rest) or copying none of the address space as in vfork (the paging tables are shared between parent and child until exec) if we can identify what memory the child will need to access before calling _exit or exec (say, "the current function and its local variables") then we could create an address space with just a few paging tables entries.
Kind of like the "zygote" forking model (early in the main process lifetime, a zygote process gets forked off, and when the main process wants another worker it asks the zygote to fork one off) except that the "zygote" is more like an induced pluripotent stem cell, having been reverted from an adult state.
interestingly enough, i thought of the same concept, except i did not get to implement that (for a few reasons). is "zygote" a term you made up or is it an established pattern?
(I don't think the Chromium developers invented it either, it's just a convenient reference.)
I think the Chrome team overthought this. If you update firefox and try to perform an action which spawns a new process, it just politely demands the user restart the browser.
(The other option would be to convince Linux distributions to implement special updater, but I am sure implementing zygote thong was easier)
Also, there's no way that libc people would want to work with the compiler people to locate the current stack frame to copy. So you'd end up with an assembly shim with a definite stack size anyway.
I mean, a good deal of the parameters you need for the relevant syscalls are strings, which means it's not sufficient to copy just the stack frame, but all the memory reachable from the stack frame. Which is a nontrivial problem if you're assuming C/C++-style code.
Described in these nicely written comments by others:
https://news.ycombinator.com/item?id=32794270 https://news.ycombinator.com/item?id=30510318 https://news.ycombinator.com/item?id=29697645
I think the best solution would be if every relevant syscall took a process handle, so you can run it either in the current process or in a non-started child process
That's not going to happen on Linux because it would be a radical change to the Linux syscall API. But if one were designing an OS from scratch today I think it would make sense to do things that way.
So if fork/clone is followed immediately by exec/execve/etc., there is minimal copying.
The implications of this is that even if you immediately execve in the child, you still have to pay for the cost of setting COW on the entire address space and then later faulting on every single writable page in the parent process. The performance impact might not be massive, but it's not nothing.
One of these enhanced fork exec calls stops the parent until the child execs. Then you don't need to touch the parent page mappings or worry about concurrency. (Although it's not ideal if the parent is threaded)
But it's been a while....
However, maybe I'm missing something, but it seems like linux already has functionality that could make spawning a process a lot more efficient and threadsafe. My idea is basically to use clone or clone3 to create a new process in a new thread group that shares the original processes memory (that is with CLONE_VM but not CLONE_THREAD). And pass a function point to call (instead of returning on the child process) and a heap-allocated stack for the child process to use.
Then there is no need to copy the address space, and you can do more things to prep before calling exec, since other threads can still release locks, you can write to memory, etc.
The downsides I see are that you wouldn't be able to safely modify the current environment variables since that would impact the parent process, and there might be some weirdness with the child process having copies of file descriptors instead of the originals. The first is easy to work around though, and the latter probably wouldn't be an issue in most cases.
Another thought I've had is that if there was a more efficient single syscall for spawning a process that combined fork and exec, even if it is a lot less flexible than fork/exec or the io_uring equivalent, something simple could probably meet the needs of most applications and benefit performance and safety in the common case where you don't need complex setup before calling execve.
Possibly just becaus clone is a linux specific API, whereas fork/exec is more portable.
In Linux you can do that with standard system calls, by spawning a thread with pthread_create() then calling vfork() from the thread. vfork() pauses only the parent thread, not the entire parent process, until the vfork child calls execve().The effect is to create a child task which has CLONE_VM but not CLONE_THREAD, which runs concurrently with all the other threads.
1. Everything needed in the "do stuff" part was known prior to the call to fork
2. Any failures in the "do stuff" part would scrap the child process and report an error to the parent process
This seems like yet another way for ferrying code/state machines into the kernel. We already have bpf.