This kind of optimisation becomes more important as the parent process's address space grows e.g. a browser wanting to spawn a process per origin or even worse an in memory database spawning a tiny shell script as hook. The existing (v)fork()+exec() syscalls have to change the existing address space to copy on write which can involve shooting down TLB entries on all other CPU cores with IPIs (inter-processor interrupts) and setup new page tables for the child only to tear them down immediately and bring in a new executable. It also doesn't play nice with unified event loops implemented with the common epoll()/kqueue() etc. syscalls which is a problem for languages like go which depend on event loops for their runtimes.