Ghosts of Unix past, part 3: Unfixable designs (2010)
lwn.net
lwn.net
I think there are applications for this kind of study. It can very clearly feed back into current practices, and possibly even more formalized language syntax that can be defensive and extensible. I would also love to see if this kind of analysis bears out which aspects of various languages (and architectural OS decisions) proved to be the most robust. Like with hemaglobin: it is one of the largest and oldest genes, it is hard to break via mutatation, and is shared by every animal with oxygenated blood cells. Something was done right with that design!
Edit: waitpid is also similarly broken and unfixable for a lot of the same reasons as signals, pidfds and waitid(P_PIDFD, ...) should be replacing most uses of that as well.
Unfortunately, it seems like this idea was rejected during the introduction of pidfd.
Edit: note that pidfds holding same-uid processes have no effect (that process already belongs to your uid), and multiple processes holding a single foreign process only count once (that single foreign process is counted toward your NPROC, since it will stick around for pidfd even if its user kills it).
However if opening a handle based on pid you may want to double-check that the handle matches your expectation before using it, since that would be prone to the same race.
Once the service forks then it becomes a problem. If you use cgroups it can be solved separately with the cgroup freezer, but there are still some open issues with this in systemd: https://github.com/systemd/systemd/issues/13101
https://lwn.net/Articles/773459/
https://lwn.net/Articles/784831/
In essence it's a classic TOCTTOU.
Under normal operation (pid returned from fork, waitpid used in one spot until the child exits, no odd sigchild shenanigans), I don't think waitpid() has any races.
Am I missing anything?
If you need to multiplex a wait for multiple processes, it's safer to use epoll with several pidfds. But you still eventually may need to call waitid(-1, ...) to reap any forked processes, and there really is no good way around that.
> Unix signals
Unix signals are an asynchronous best-effort form of out-of-band IPC. Because the programs that make up your Unix application already use pipes for IPC (which are synchronous and reliable), the role of the signal handler in a program would be to either absorb the signal by taking some localized action, or translate the signal into some piped IPC message to other program(s) in the application to consume and handle.
It's been pointed out elsewhere that threads and signals don't play nicely. But that shouldn't be a problem for a multi-program Unix application -- you'd keep the multi-threaded logic in a separate program(s) from the signal-handling logic, and have the signal-handling program forward the multi-threaded program the signal data in-band, via a pipe. For example, you might factor the application into a supervisor program and one or more subordinate programs (which can be multi-threaded), and have the supervisor intercept signals and route the relevant IPC notification to subordinates via pipes.
> Unix permissions
The "one user and one group" model for files stops being so limiting if you can make it so the different programs that make up your application run as different users and groups. For example, a "logger" program in your application would have a separate user/group ID than a "database" program, and in doing so, ensure that the "logger" program can only access log state, and the "database" program can only access database state.
I don't think Unix signal behavior is the problem. The problems outlined in the article stem from people using them inappropriately. The new signal syscalls introduced in Linux over the years haven't stopped people from misusing them.
> You can't really move all signal handling logic to a dedicated process, because every process can receive signals.
Processes are not obliged to take action in response to signals. But, they could simply propagate the signal data to the parent process via a pipe file descriptor it inherits. Then, you could place all the signal-handling logic into a supervisor -- the supervisor would get notified via a pipe when one of its descendants receives a signal, and take appropriate action.
> And as for databases: every major production database I've seen implements its own authentication and permission scheme, so it can do things like provide ACLs on a more granular (per-table, sometimes per-row) basis.
No one said an application can't have its own authentication and permission scheme. All I am saying is that if you factor your application into multiple processes running under different system-level user accounts, you can get more mileage out of the Unix permission system than you could otherwise, because the kernel would be able to distinguish individual pieces of your application as having different sets of permissions.
It seems pretty clear that they're used for way more than originally expected (did threads even exist when signals began?) - and I suspect a number of systems use other communication paths already.
The question is, was that late? Windows NT had kernel threads from the beginning, so maybe a few years earlier. But then it took NT years to become stable enough to be used in servers, so saying they were generally ahead would not be a correct description.
So if Unix is considered late (according to the GGP) and NT not a comparable competitor, who was really early? If anybody.
[1] https://www.cs.utexas.edu/users/EWD/transcriptions/EWD13xx/E...
It is not as if the UNIX culture was to pay attention to best practices being done on other systems.
So it used very limited tools, like C with its very simple semantics and a preprocessor instead of a module system, very simple kernel mechanisms, etc.
Had AT&T been allowed to sell it, I bet we wouldn't have turned with it everywhere.
Signals themselves got new APIs long before signalfd: sigaction and posix real-time signals were already a thing, as were posix threads, when Linux was invented.
This is reasonable in a memory-safe language, but means you need some other way to interrupt a read from stdin with ^C.
Commands providing "why user X can't access Y" and recommended solutions can help.
The POSIX draft ACLs had the same problem, where a chmod might not grant you the permission that you're asking for. Back when Solaris implemented POSIX draft ACLs, they needed to change many user-level interfaces (e.g., the chmod command and the ftp daemon) to have a chmod request work the way end users expected.
Happy to argue it or simply be told I'm wrong, but I've yet to encounter a not-insane permissions model that I couldn't solve with some "simple" nested groups (that in and of itself is a tooling problem, but a solvable one) and POSIX.
I spent some... long nights writing a program that ensured ACLs were appropriately applied in complex setup and automatically inherited by new files.
It supported both NFSv4/ZFS format and POSIX ACLs... the former was essentially one line per ACL, the latters involved very racy "drop all", "apply again recursively in very specific order" for every change in ACLs.
Speaking of signals, the mess of when POSIX threads collided with signals. And with fork. And chdir being process wide ...
On a more meta note, open source means a lot of different things. There is actually a lot of nuance in the different styles of open source. Linux vs chromium vs that project that just does source dumps. Whether they accept contributions, accept bug reports/feature requests, allow you to build from source (source dumps often don't include a working build system), have open communication channels, etc can all vary. I hope we have more specific terms for the different styles of open source in the future.
Project leadership is the most significant aspect - look at Linus's absolute declaration that the kernel can "never break userspace" meaning that once an API is exported to userspace it never gets removed.
This is actually similar to Microsoft's philosophy though theirs is more "business oriented" (nobody will buy Win95 if their DOS and Win 3.1 programs won't work). Another example of this is Knuth's TeX code.
Open source developments seem to lean (in general) more toward "rip it out and replace everything" (see for example internal kernel APIs not exported to userspace) because access to the source means they can fix the things that touch it. Closed source programs more likely just die and get entirely replaced, otherwise they roughly try to keep working as is.
It seems like something that should be so simple, but once you sit down and try to build it you'll realize you have to support so many uses cases. I bet if you asked everyone on HN how they'd do it, you'd end up with so many confident answers that also had shortcomings themselves.