Why threads can't fork
thorstenball.com
thorstenball.com
UNIX "fork" started as a hack. The reason UNIX originally used a fork/exec approach to program launch was to conserve memory on the PDP-11. "fork" originally worked by swapping the process out to disk. Then, at the moment there was a good copy in both memory and on disk, the process table entry was duplicated, with one copy pointing to the swapped-out process image and one pointed to the in-memory copy. The regular swapping system took it from there.
Then, as machines got bigger, the Berkeley BSD crowd, rather than simply introducing a "run" primitive, hacked on a number of variants of "fork" to make program launch more efficient. That's what we mostly have today. Plan 9 supported more variants in a more rational way; you could choose whether to share or copy code, data, stack, and I think file opens. The Plan 9 paper says that most of the variants proved useful. But that approach ended with Plan 9. UCLA Locus just had a "run" primitive; Locus ran on multiple machines without shared memory, so "fork" wasn't too helpful there.
Threads came long after "fork" in the UNIX world. (Threads originated with UNIVAC 1108 Exec 8 (called OS 2200 today), first released in 1967, where they were called "activities"). Exec 8 had "run", and an activity could fork off another activity, but not process-level fork. Activities were available both inside and outside the OS kernel, decades before UNIX.
That's why UNIX thread semantics have always been troublesome, especially where signals are involved. They were added on, not designed in.
http://cm.bell-labs.com/cm/cs/who/dmr/man21.pdf http://www3.alcatel-lucent.com/bstj/vol57-1978/articles/bstj...
John Lion's "A commentary on the Sixth Edition UNIX Operating System", goes through the PDP-11 UNIX kernel line by line.
http://www.lemis.com/grog/Documentation/Lions/book.pdf
Page 37, Section 7.13, "newproc":
1906: Call "xswap" (4368) to copy the data segments into the disk swap area. Because the second parameter is zero, the main memory area will not be released.
1907: Mark the new process as "swapped out".
1908: Return the current process to its normal state.
Anyhow, there's a difference between the fork interface being a hack, and the fork implementation being a hack. Unix is is a cornucopia of implementation hacks. That doesn't mean the interfaces weren't deliberately and thoughtfully designed.
Much like C, what makes Unix unique and still relevant is that the deliberate design took into account practical implementation considerations. Unix and C are most elegant from an engineer's perspective. It's an interesting balance of interface complexity and implementation complexity. This is why some people claim that the Unix design philosophy epitomizes "worse is better".
Sounds like linux's "clone" system call. Which is the underlying syscall which clib's fork() uses.
You can do just about anything imaginable with it: http://linux.die.net/man/2/clone
For example: you could create a child process-like-thing which shares nothing but the signal handler table. No idea what that would be good for.
There's also the matter of API complexity. If you want to configure the context in which a program runs, there are roughly two different API approaches you can take. You can deeply parameterize the 'run' command, or you can configure the "current" process and then hand that off to the new process. But you usually need the second API (configuring the current process) anyway.
So I don't see fork(), per se, as a particularly crufty API. Slightly more configurability, like clone() or Plan 9 is better.
Multi-threaded signals are pretty dreadful, though. In a multi-threaded system, signals should be delivered on a separate thread, not via an existing thread getting hijacked.
That is basically the recommended way to handle signals in a pthreaded program - have one dedicated signal-handling thread that calls `sigwait()` in a loop, and block all signals in the signal masks of the other threads.
There have been discussions about this in OS design for almost forever.
An example of a problem we encountered: what if execve() fails? Then we want to print an error message, based on the value of errno. perror() is the preferred way to do this. But on Linux, perror() calls into gettext to show a localized message, and gettext takes a lock, and then you deadlock.
This is a hard problem to solve because it requires knowing the status of all locks. This breaks the abstractions that library authors present, where locks are internal and not exposed.
If Go figured out how to efficiently detect data dependencies and relationships, and then automatically move goroutines around, that would be exceptionally noteworthy. Everybody starts out thinking they can do this, which is why Solaris, NetBSD, Linux, Java, et al all started with out M:N threading. But then when they figure out that it's a Really Hard(tm) problem, they invariably shift to 1:1.
I've found that it's better to leave it to the developer to choose whether to run an OS thread or coroutine, just as the developer chooses between a process and OS thread. So in my project[1] I don't spend much time trying to automate that.
[1] http://25thandclement.com/~william/projects/cqueues.html
The order of calls to pthread_atfork() is significant. The parent and child fork handlers shall be called in the order in which they were established by calls to pthread_atfork(). The prepare fork handlers shall be called in the opposite order.
I'm not sure if that's the best approach, but it's an attempt at least.
It has a few corner cases across POSIX systems. I wouldn't bet it works 100% the same way in all UNIXes.
I used to do cross platform across Aix, HP-UX, Solaris, GNU/Linux, FreeBSD and Windows NT/2000 back in the .COM days.
Excepting Windows, I rarely run into difficult portability problems except when I deliberately use non-POSIX functionality or newer POSIX functionality.
I target Linux, OS X, OpenBSD, NetBSD, FreeBSD, Solaris, and AIX. The biggest laggard was OpenBSD, particularly wrt to threading, signal handling, and real-time extensions. But in the past couple of years that's been substantially addressed.
One of my biggest headaches now is OS X. They appear to have stopped trying to track POSIX, so while everybody else is busily implementing POSIX-2004, POSIX-2008, and tentative POSIX features, OS X is nearly at a stand-still. OS X hasn't fixed any significant conformance issues, adopted real-time extensions, nor adopted any POSIX-2008 features for several years, now.
Was this a common (no pun intended) problem among CL implementations and why server daemonization is an issue? I am just learning CL, and noticed RESTAS only really supports this daemonization feature in SBCL, if I believe the documentation.
I guess I am going to have to dive into the source this weekend and check it out.
Judging from the one sentence on the landing page, I supposed it would blow up. Haha.
Also, just imagine something like this:
file.write(blah);
if (has_forked) thread_exit();
file.write(blah2);
if (has_forked) thread_exit();
file.write(blah3)
if (has_forked) thread_exit();
Note the race conditions in the above code... forkall() is a disaster. Presumably you'd combine it with pthread_atfork(), but it's still a total mess.fork() is one of those things that's conceptually quite elegant but in practice has too many problematic edge cases. (Signals seem to fall into the same group.)
I do like that fork, then setup functions, then exec keeps the code simpler. If everything is done as a single call like CreateProcess then it needs to have a way of passing what the setup functions need as data - rather a lot of it.
Meanwhile the Linux/Unix decision was to have about a dozen different APIs to do variants of the same thing (fork, vfork, and clone plus execl, execlp, execle, execv, execvp, execvpe, fexecve, posix_spawn, posix_spawnp, and probably more). This isn't really any less complex.
The argument complexity of CreateProcess could not possibly derive from merging fork() and exec() because fork() has no arguments.
fork(); open(); close(); dup2(); fdcntl(); sigprocmask(); sigaction(); setrlimit(); prctl(); chdir(); chroot(); capset(); execve();
(and perhaps more besides, with new ones potentially added with new features, for example seccomp();)
The question is whether it's better (more elegant) to load that complexity onto one complex function call, or farm it out to multiple simple function calls. It is apparent that this is a question in which personal taste is involved.
The brilliance of fork/exec is in the realization that the tasks you do during process creation are the same as the tasks you do during normal execution, and therefore you can just re-use those functions. For example, CreateProcess has a parameter for the current directory; fork/exec doesn't need one because you just use chdir on the child side of fork.
The proliferation of exec* is indeed baffling, but these are just dumb variants trying to make the exec call itself "easier." They don't reflect any essential complexity.
If you need to use a language with a runtime (not just Go by any means, the likes of Python also suffer from this issue) from two processes that need to be separate but communicate with each other, do the fork first, then start the language runtime (i.e. embed the language in your parent program).
You can try to bundle it all up into a single API like posix_spawn, but the API becomes large and it's hard to cover everything. posix_spawn sure doesn't.
fork is an elegant solution to this problem: you can run code as the child, before the child gets to start. I am not aware of any alternative that's as flexible.
fork() allows you to set up file descriptors such as stdin and stdout before execing another process. This is essential for pipelines.
Android uses this to pre-load framework resources and code in a way that lets all applications share the backing memory. And when applications crash, they don't bring down the initial process that preloaded everything (so it can continue spawning new apps).
How would you handle that with only threads or spawning another process?
> Better to only share memory that processes explicitly want to share, and make it clear which one owns any given region of memory.
We are only sharing memory that we explicitly want to share: we load only what we care about, then fork.
The proof that the problem vs. forking is not threads, is that the worse example given in the article referenced, that of file I/O, occurs as well in processes without threads (or "single-thread" processes if you want): if you write to a file from both the parent and the child, you must take precautions at the application level. This has nothing to do with threads.
The article has perhaps the most likely example. Imagine malloc() has a lock. Another thread happens to be inside malloc() at the moment that you fork(), and therefore owns the lock and might be in the middle of manipulating shared data structures. Now suddenly the child cannot malloc(), because that thread [suspended in the middle of its execution] isn't going to be carried over to it, and will never be able to clean up its intermediate state and release the lock.
Edit: IMO it's better to just admit the programming model is thorny and move on. I feel similarly about signals. Restrict what you do after fork() and generally be cautious, the same way you'd be cautious reacting to a signal.
Otherwise you wouldn't offer locks in languages either!
This is a feature that can be implemented.
I used to do cross platform across Aix, HP-UX, Solaris, GNU/Linux, FreeBSD and Windows NT/2000 back in the .COM days.
See also: high performance event handling, etc. But does anyone suggest you just can't ever use lots of sockets? (It is a bad idea sometimes, to be fair!)
We write the code necessary to do it.
(Since we're apparently going into personal history, I used to be a platform manager for a ~million line cross platform codebase that ran on most of the things in your list plus IRIX and OS X. Most of my responsibilities were really about keeping the build environment working enough to run one binary across all the Linux distributions, but I was also involved in issues that specific to my platform. That is one of many projects where I've dealt with platform specific issues on a variety of codebases, several of which I've written in their entirety.
Anyway.
I have a passing familiarity with some of these issues.)
I didn't want to do some kind of credential call, it was more to explain where I got my experience from, as many in forums tend to think all UNIXes are 100% alike, or worse, GNU/LINUX === UNIX.