io_uring_spawn: Launching New Processes with io_uring
phoronix.com
phoronix.com
I look forward to seeing the impact that this change has, for example, on future builds. A significant amount of time is spent at the beginning of a process spawn, setting things up - being able to batch things up without the process startup load, is going to make things much, much faster. Process pools? Yes please! :) I will definitely be a customer of posix_spawn() when its ready ..
One has to wonder, though, given Josh' work with Rust, if any of this work is going to be done in Rust? One can dream ..
Not yet, since the Rust-in-Linux work hasn't been merged yet. And even once it is, it won't be usable in the core kernel yet. But I'm hopeful that io_uring will be able to make use of Rust in the future.
And the userspace bits will definitely be possible to do in Rust. As soon as this ships in a stable kernel, I plan to add a code path in the Rust standard library to launch processes this way. :)
edit: K I've reached my HN rate limit again, thanks Dang.
Anyway, here's a source for the Docker seccomp filter since someone replied inquiring.
https://github.com/thaJeztah/docker/blob/master/profiles/sec...
also, iouring doesn't allow arbitrary syscalls. so it's not really a complete seccomp bypass even when allowed
And the need to implement all the functionality twice. Who is going to do that?
I guess shells (eg with xargs) would be one example.
Edit: the talk mentions build systems as a motivating use case.
What I missed from the presentation is some discussion about the asyncronicity of the clone calls. Yes, the average time of cloning once is 30us, but if you queue up two of them at once, how long does that take? Will it just sum or will the total time be shorter? Today there are applications that set up a thread pool by recursively open new threads, both in the original thread and in the forked ones. That is more than a little bit icky and could perhaps be simplified with this.
I'm working on a new architecture for this right now that should make this case work well.
(Note that threads on Linux are just processes which share memory and some other properties.)
Merely halving it may not be all that's needed to totally replace those runtimes but it's a good step in the right direction.
The cost of fork is proportional to the complexity of the spawning process (e.g., the size of its page tables), not the spawned process. The talk gives an example where fork takes 7ms, although I'm not entirely clear what it's doing (I only read the slides, not listened to the talk).
vfork and posix_spawn allow you to avoid that complexity. However, this comes at a price: posix_spawn only permits a very limited set of process attribute twiddling, and vfork is so unsafe that it's difficult to sanction its use... well in general.
From the slides, the intention is to provide something like vfork except without the inherent unsafety of vfork. It describes vfork as "Effectively a thread with no synchronization running with the same stack as the parent". Meanwhile, you can look at io_uring as describing a kernel thread that does a limited set of tasks... so why not implement something like vfork on top of io_uring? Of course, the interesting question is what process bits can be twiddled in io_uring_spawn, but unfortunately, there's no examples from what I can see to get a sense of how powerful it is or isn't.
The last benchmark column touches 1gb of memory,while the second just allocates without modifying the memory, so I guess all that overhead is page table CoW mapping.
During QnA they also talk about exposing as many config calls as possible via uring.
I allocated 1GB of memory, wrote a byte to every 4KB page of that memory, then did the same fork benchmark. The kernel has to do much more work to copy the page metadata associated with the resulting 262144 page frames.
> Meanwhile, you can look at io_uring as describing a kernel thread that does a limited set of tasks... so why not implement something like vfork on top of io_uring? Of course, the interesting question is what process bits can be twiddled in io_uring_spawn, but unfortunately, there's no examples from what I can see to get a sense of how powerful it is or isn't.
We're going to need to add a number of additional io_uring operations to provide the full capabilities of posix_spawn, but I'm expecting to do so incrementally. Eventually: everything you can do by making a syscall, at a minimum.
Since it pauses the execution of the parent process (or at least the parent thread?) it gives a window for the child to use the current stack frame to set up an exec() call.
When I first learned about vfork() (early 2000s) I had the impression it had become a bit pointless since CoW sharing had made fork() really fast. But with modern-sized address spaces it is (has become?) horribly slow again.
Also, processes are an awesome abstraction - they give security and reliability. Reducing their cost is nice.
With a huge process, you have a timeframe between the child is spawned and it executes exec*() where you typically "do stuff" (such as closing a lot of fd)
During this timeframe the parent process has its universe COW'ed, and each write will trigger a page fault.
The performance impact can be concerning in the _parent_ process.
Now fork+exec isn't the biggest chunk of that. But it fingers the memory subsystem. On many-core systems this leads to lock contention. So the cheaper starting a new process gets in terms of virtual memory operations the less lock-contention you get overall which also speeds up other parts.
I agree that's a useful comparison, although the status quo is also worth benchmarking. A fork_exec syscall sounds like a posix_spawn syscall. IIRC, glibc's current posix_spawn is a library function that does a vfork followed by setup operations that portable code isn't allowed to call after vfork.
I think there are two dimensions in which io_uring can be used to amortize/eliminate syscalls:
* sequence of operations: in this case, (v)fork, then setup calls, then exec. This sequence is unusually expensive because doing them separately requires extra page table work (particularly for the fork case, but even for vfork vs having a blank process). A posix_spawn syscall could achieve similar efficiency gains as this io_uring_spawn.
* batches of operations: in this case, lots of processes to spawn. It sounds like he hasn't posted numbers on that and intends to: "He's looking forward to ... supporting a pre-spawned process pool". If that turns out well, then the io_uring approach seems worth doing, and also can be used to implement a new in-user-space posix_spawn, for less total syscall surface area than doing both io_uring and a synchronous posix_spawn syscall. Maybe the supported setup operations also can be in common with io_uring operations for other purposes. I'm curious about that but don't see a link to benchmark code / implementation code / manpage.
1. posix_spawn() performance seems terrible on Linux. There is just no excuse for it to be worse than vfork()... I know on Linux it is implemented in library code, but on macOS it is a syscall, which achieves most of the stated benefits of io_uring_spawn (batching up all the posix_spawn_actions and then handing them down to the kennel in a single syscall).
2. Again, I think the whole "hope posix_spawn supports the actions you need" is a bit overblown here. It is not like io_uring supports operations for every single syscall that exists either. People keep holding up the lack of some operation in posix_spawn() as a justification for designing a new interface. Just add some new flags and actions (see posix_spawn_file_actions_addchdir_np(), POSIX_SPAWN_SETEXEC, or POSIX_SPAWN_CLOEXEC_DEFAULT on macOS for examples).
Obviously posix_spawn() is a synchronous operation, so in that respect io_uring_spawn is inherently a bit faster for processes that want to asynchronously launch processes, but if Linux implemented posix_spawn() like macOS did then you could achieve the same thing just by adding support for issuing a posix_spawn() operation via io_uring just like any other operation.
Given that io_uring has a lot of momentum and is getting broad support for various operations it probably does make sense to implement io_uring_spawn and then layer posix_spawn() on top of it as proposed in the talk, but I wish the slides called out that a lot of performance gains seem (at least me) to have little to do with the nature of io_uring and more to do with that fact that posix_spawn() is implemented purely in userspace on Linux and can't batch the existing actions because of that.
But also, the tricks to make it faster without using io_uring are much less safe than using io_uring. vfork is dangerous.
io_uring is the standard kernel mechanism for doing multiple operations in one syscall, so it seems like the most logical fit for this.
Jens Axboe (io_uring maintainer) and I have many more plans to make this faster.
Leaving that aside, there are many additional capabilities this unlocks. For instance, we can maintain a pool of processes set up and ready to exec.
I think there are at least some niche cases in which folks would still want it on. For example, some processes may rely on anonymous mappings being cheap until first write via Linux's "empty_zero_page". IIRC, they once tried to remove this concept but put it back in. Some context here: https://lwn.net/Articles/340370/
Still need resources behind the pit stop, lots and lots of resources.