I'd say this is depending on perspective both true and false¹, but also unhelpful to work with here.
Instead, I would suggest this perspective: the kernel has neither processes nor threads; it has tasks, which are entities the scheduler can run. They're exposed to userland as processes and threads. Excluding kernel tasks/threads, which can have arbitrary rules but are also user-visible, a task is exposed as a thread, and a set of threads is exposed as a process. Both operations working with threads as well as operations working with processes exist.
We're looking at an API in this case that works with threads on one side (parent, the signal is triggered by thread exit) and processes on the other (child, the signal is process-targeted). How these were created is irrelevant, what matters is the abstractions they refer to.
¹ you could equally well argue that processes do not exist in the kernel, they're just threads with sharing flags set differently.
If you want to register per-thread signal handlers you're forced to step outside the bounds of glibc and pthreads which I think is quite unfortunate.
Now, pthread_t is actually a pointer to an undocumented structure, and the TID is stored at a certain offset in it… so it is easy to pull the TID from there. Until some day the glibc developers change the layout of the structure and suddenly that code breaks.
There’s an entry in glibc’s bug tracker for this - https://sourceware.org/bugzilla/show_bug.cgi?id=27880 - but it doesn’t look like it will be implemented any time soon
Digressing the conversation further, I notice in the docs that CLONE_SIGHAND requires CLONE_VM and CLONE_THREAD requires CLONE_SIGHAND. Any idea if there's a technical reason for that or is it just POSIX constraints needlessly infecting the kernel?
It's particularly confusing that the TID is the real identifier but the documentation generally refers to scheduling entities as processes. So you use a TID to refer to a process and a PID to refer to a thread group ... right. Very straightforward.
pthreads (POSIX threads) is itself a POSIX API
I guess one reason why it doesn’t have any TID concept, is although Linux nowadays uses 1:1 threading (one kernel thread per user-space thread), historically many Unix thread libraries were designed to use 1:N threading (a single kernel thread runs multiple user space threads) or M:N threading (a pool of kernel threads runs a pool of user space threads where the two pools differ in size). Plus, while Linux went with the model that processes and threads are basically two slightly different variants of the same thing, in other POSIX implementations they are completely distinct object types. Since pthreads are designed to support such a wide variety of implementation strategies, they can’t assume threads have any kernel-maintained unique ID, because in some of those implementation strategies there might not be one.
> I notice in the docs that CLONE_SIGHAND requires CLONE_VM
I think this is necessary? If it wasn’t, the child might load new code (dlopen or JIT) and then install a signal handler pointing to it. With CLONE_SIGHAND, it shares signal handler with parent. But without CLONE_VM, the memory mapping containing the new code wouldn’t exist in the parent, meaning instant segfault as soon as the signal is delivered
> and CLONE_THREAD requires CLONE_SIGHAND. Any idea if there's a technical reason for that or is it just POSIX constraints needlessly infecting the kernel?
Well, this one is more POSIX (and historical Unix before it). Signals are primarily a process-level construct in Unix/POSIX, not thread-level – since back when signals were invented, threads hadn’t been invented yet (on Unix–PL/I running under OS/360 MVT already had multithreading, which it called 'multitasking', in 1968, Unix development didn't start until 1969). Although we’ve now got per-thread signal masks and thread-directed signals, the actual handlers are still per-process
Also, I think another reason for disallowing certain combination of clone() flags: obscure combinations can expose bugs, possibly even security vulnerabilities; if there is no great demand for a specific combination, the kernel devs may conclude it is safest to disallow it