The history of sending signals to Unix process groups
utcc.utoronto.ca
utcc.utoronto.ca
In modern times wer'e doing more thing like containerizing a bunch of processes they are weakly related instead of using a VM for that so the use case became more pressing. There is cgroups for managing such a group of processes. I'm suprised to learn there is no way to reliably send a signal to all processes in a cgroup though.
> SIGKILL will actually leave lots of orphans around instead of cleaning everything.
Well actually orphan processes exist for a reason. Your supposed to "wait" / reap them to make them non orphan.
I know not being a library has different considerations, but some ideas I used in timeout(1) to kill a process group may be useful. Tricky things like using sigsuspend() to avoid signal handling races.
https://github.com/coreutils/coreutils/blob/master/src/timeo...
cgroups might be another avenue to explore, being more modern and so having less compatibility baggage
You need to use a kernel feature intended for this purpose, which is what things like process groups (linked article), jobs (80's BSD feature) and cgroups (modern generalization of the idea) were designed for.
>Linux is in the middle of adding new APIs like pidfd_send_signal, but none of them are aimed at improving the situation with grandchildren.
While my subreaper is vulnerable to pid reuse, but I think it could be fixed by having it do this:
loop
children = scan-for-children
for each child
if no pidfd for given child
pidfd open child ( dropping if error )
still-children = scan-for-children
close-and-drop any pidfd's that are no longer children
safely-kill-children-via-pidfds
If you kill the direct children, the grandchildren will become direct children since you are a subreaper, and then so on.In summary, a helper process using subreaper+pidfd should be able to properly and safely contain and kill grandchild processes.
So if you want to write software that can target MacOS or any of the BSDs, then you can't use cgroups.
Maybe Linux should implement something like pidfd_getpfd() (returning the PID FD of the parent process) and pidfd_fork() (returning the PID FD of the child process).
There may be a better way to do this, but one way might be to open() /proc/<PID> and use the resulting FD to both to check /proc/<PID>/cgroup (using openat(FD, "cgroup", O_RDONLY|O_NOCTTY)) and then use the same FD as a PIDFD when calling pidfd_send_signal().
ISTM that systemd should use this method if its current method is just looping over PIDs.
Edit: Indeed both your and my methods were proposed to systemd: https://github.com/systemd/systemd/issues/13101 It's just waiting for someone to implement it.
This, IIRC used to be the standard behavior back in the days, but in recent years, for some reasons, things have not been so simple.
Looking at the man page of ps(1) sort of starts to elucidate the effing mess that the name "group" summons the in unix process world.
Are we talking about the real group id (RGID), the effective group id (EGID), the controlling progress group id (TPGID), the control group (CGROUP) the textual group id (EGROUP), the filesystem group id (FGID), the textual filesystem group id (FGID), the process id of the process group leader (PGID), the saved group id (SGID), the session id (SID), the supplementary group id (SUPGID), the thread group id (TGID) ?
I'm pretty sure I'm missing some.
What an effing mess.
Catching SIGINT and detaching from the controlling terminal are pretty old concepts.
are you sure? let's assume you typed "service httpd start". it starts the init script which starts the httpd in the background. IMO most in use-cases you don't want to kill everything ever forked from the current shell.
* http://jdebp.info./Softwares/nosh/bsd-service-command.html
Things that this history does not cover include System 5 Release 2 "shell layers". Shell layers multiplexed virtual terminals onto a single physical terminal, preceding its obvious successors screen by somewhere around three years, and tmux by about quarter of a century. The popularity of tmux and screen nowadays shows that the BSD job control mechanism did not entirely kill off the vision of shell layers, as I recall people thinking it had at the time.
In the past, people have popped up to say “but that’s only in the bind() implementation side” (when talking about `sockaddr`, for instance) but I find this argument amusing. At the end of the day, if something breaks because of this, that nuance is lost amidst the mass of frustration.
People often point to the function calling conventions of the API, such as the "sockaddr" structures and the "errno" mechanism (that has historically been tricky for non-Unix operating systems and for systems where there are multiple C compilers with multiple C runtime libraries). But there are actual architectural decisions that had fairly reasonable alternative paths. I wouldn't say that it's a "mess" because of these alternatives, but it is definitely not the sole and unequivocal way of approaching things; and there were and are tradeoffs to be had.
Consider an alternative design (Plan 9 is something like this, it's been a while so I'm short on details):
1. Open /net/dns and write the desired hostname, then read back the resolved IP address
2. Open /net/tcp/clone and write the address and port, then read back a connection number
3. Open /net/tcp/$conn_id and the file is now a full duplex TCP stream
Compare this to BSD sockets: use getaddrinfo for the DNS lookup (gross!), create a socket with a syscall, then connect the socket with another syscall using sockaddr (gross!). Much worse and much less Unixy.
[1] that's the ideal of course, reality is different.
All the “it’s all a file” idea was added as an ideal along the way because the interface was so dang convenient. But I’ve never read it in any documentation.
Even pipes were added as a “that’s a cool idea.” There was a conversation, a late night coding session and pipes came into existence.
Plan 9 was a response to the evolved (vs designed) Unix. Took “the best” from it and made those design principles. It was an “if we could do it all over again, knowing what we know now” project.
What’s interesting to me is that evolved systems seem to dominate design-first systems in adoption. Maybe it’s that they’re pragmatic. I don’t know. Or maybe my observation is just wrong.
If you want to establish bi-directional communication with some process on the same host, that process should create two rendezvous points with mkfifo(2). Your process opens the read FIFO for reading, and the write FIFO for writing, and you're done. If you finish writing and want to just read, you close(2) the write file descriptor, and keep reading the read file-descriptor until you hit EOF.
If you want to establish bi-directional communication with some process on a remote host over TCP/IP, that process needs to listen(2). Your process connects to it with connect(2), and gets back a single-file descriptor that (unlike any other file-descriptor) is both writable and readable. If you finish writing and want to just read, you have to use the special shutdown(2) function, a quirk that only exist to work around the quirk of TCP sockets being both readable and writable.
Some might argue that this is a pretty minor wart, all things considered, and sure, it's nowhere near as confusing as the mess of terminal job control. But it also seems like it they could have implemented it in a non-quirky way if they'd just spent thirty seconds thinking about it beforehand.
Processor affinity, sequential execution on available/shared resources, hold queues for non-fatal errors, and policy enforcement are missing from job control.
This is how they optimised it in the days before COW, by fully preserving the intended behavioural semantics of fork/exec. You shouldn't modify memory in the child and the OS therefore shouldn't copy the memory space.
Essentially you run into the same problem - now the complexity of the spawn(...) interface is shifted to that binary.
You could simply use ordinary IPC for these things, though. They need not be implemented as part of the OS.
One example among many warts: in traditional Unix, user programs manipulate process IDs in a shared PID address space, not process handles private to one process (as in file descriptors). Therefore, nearly any operation done on a process through a PID is inherently racy. That's because between acquiring the PID of a process you want to interact with and performing an operation by PID, the PID could no longer refer to the process you were interested in (for example, the process could die and the kernel could recycle the PID for a new, unrelated process). There are ways to fix that (pidfd_open on Linux for example), but the sheer amount of legacy code out there means that these warts will stay around in Unix-like systems for a very, very long time.
Even if you somehow fix every single design issue with Unix and remove all the warts, the end result would not really look like Unix anymore. One example of a legacy-free, capability-based operating system with no implicit ambient authority is Fuchsia, whose kernel interfaces do not resemble traditional Unix syscalls at all.
The concept has been revived (in a notoriously clunky way that's likely not directly compatible with previous facilities) via multiseat support in more recent Linux distributions.
The example I'm most familiar with, because I work on it, is Shadow. We used ptrace for a bit but now use seccomp.