I've taken to calling such a process a "lich," because it's not quite a zombie - the parent can reap a zombie by calling wait(), but the lich has used magic to avoid true death.
I've taken to calling such a process a "lich," because it's not quite a zombie - the parent can reap a zombie by calling wait(), but the lich has used magic to avoid true death.
But TASK_KILLABLE is not used in most places it should be. Patching your (least) favorite driver to use TASK_KILLABLE could be a good entry point to contributing to the kernel.
This older, tangential HN discussion and the comments on LWN have a bit more info: https://news.ycombinator.com/item?id=18056946
That of course doesn't work if you're dealing with a filesystem that still uses TASK_UNINTERRUPTIBLE
[1] https://docs.microsoft.com/en-us/windows-hardware/drivers/dd...
[2] https://docs.microsoft.com/en-us/windows-hardware/drivers/ke...
It was a java program that I killed -9. It died but was never reaped.
It prevented a software shutdown so I had to use the ol' STOP-A
In the end, you really have little choice but to trust what the OS is telling you, and if the OS is lying for whatever reason, to fix the OS. As hard as that last bit may be as a solution, it still turns out to be the easiest, most-effective fix. And it may not be a perfect fix, either... but it'll still be the best.
That doesn't solve the kernel thread being stuck forever problem. I'm not sure what the fix is there... is it bad drivers with no timeout mechanism? I don't know how they'd do that unless the kernel IO is all async anyway. Just tearing down the kernel thread seems likely to leave the locks abandoned in bad states.
Sounds like a hack.
So while I agree with the other posts that a supervisor can't recover from this state, it can be aware that it's happened. And it's quite common for supervisors to check for readiness and liveness beyond "well the OS says it's running, sounds good to me."
That's why a process that wants to stay running has to provide some form of heartbeat API.
Can a malicious process use this to avoid being killed
A malicious process can't do useful work - as soon as the kernel gets unstuck, the signal will get delivered to the thread and kill it.