I guess shells (eg with xargs) would be one example.
Edit: the talk mentions build systems as a motivating use case.
I guess shells (eg with xargs) would be one example.
Edit: the talk mentions build systems as a motivating use case.
The cost of fork is proportional to the complexity of the spawning process (e.g., the size of its page tables), not the spawned process. The talk gives an example where fork takes 7ms, although I'm not entirely clear what it's doing (I only read the slides, not listened to the talk).
vfork and posix_spawn allow you to avoid that complexity. However, this comes at a price: posix_spawn only permits a very limited set of process attribute twiddling, and vfork is so unsafe that it's difficult to sanction its use... well in general.
From the slides, the intention is to provide something like vfork except without the inherent unsafety of vfork. It describes vfork as "Effectively a thread with no synchronization running with the same stack as the parent". Meanwhile, you can look at io_uring as describing a kernel thread that does a limited set of tasks... so why not implement something like vfork on top of io_uring? Of course, the interesting question is what process bits can be twiddled in io_uring_spawn, but unfortunately, there's no examples from what I can see to get a sense of how powerful it is or isn't.
The last benchmark column touches 1gb of memory,while the second just allocates without modifying the memory, so I guess all that overhead is page table CoW mapping.
During QnA they also talk about exposing as many config calls as possible via uring.
I allocated 1GB of memory, wrote a byte to every 4KB page of that memory, then did the same fork benchmark. The kernel has to do much more work to copy the page metadata associated with the resulting 262144 page frames.
> Meanwhile, you can look at io_uring as describing a kernel thread that does a limited set of tasks... so why not implement something like vfork on top of io_uring? Of course, the interesting question is what process bits can be twiddled in io_uring_spawn, but unfortunately, there's no examples from what I can see to get a sense of how powerful it is or isn't.
We're going to need to add a number of additional io_uring operations to provide the full capabilities of posix_spawn, but I'm expecting to do so incrementally. Eventually: everything you can do by making a syscall, at a minimum.
What I missed from the presentation is some discussion about the asyncronicity of the clone calls. Yes, the average time of cloning once is 30us, but if you queue up two of them at once, how long does that take? Will it just sum or will the total time be shorter? Today there are applications that set up a thread pool by recursively open new threads, both in the original thread and in the forked ones. That is more than a little bit icky and could perhaps be simplified with this.
I'm working on a new architecture for this right now that should make this case work well.
With a huge process, you have a timeframe between the child is spawned and it executes exec*() where you typically "do stuff" (such as closing a lot of fd)
During this timeframe the parent process has its universe COW'ed, and each write will trigger a page fault.
The performance impact can be concerning in the _parent_ process.
Also, processes are an awesome abstraction - they give security and reliability. Reducing their cost is nice.
(Note that threads on Linux are just processes which share memory and some other properties.)
Merely halving it may not be all that's needed to totally replace those runtimes but it's a good step in the right direction.
Since it pauses the execution of the parent process (or at least the parent thread?) it gives a window for the child to use the current stack frame to set up an exec() call.
When I first learned about vfork() (early 2000s) I had the impression it had become a bit pointless since CoW sharing had made fork() really fast. But with modern-sized address spaces it is (has become?) horribly slow again.
Now fork+exec isn't the biggest chunk of that. But it fingers the memory subsystem. On many-core systems this leads to lock contention. So the cheaper starting a new process gets in terms of virtual memory operations the less lock-contention you get overall which also speeds up other parts.