Kill Me Softly – Kill processes in a reliable manner
github.com
github.com
Handle SIGTERM only if you get a blackbox-like application (a database that will break if you pull the power plug on it) and so on.
So you have write code to handle abrupt power shutoff? And test it, and make sure it doesn't break your data at rest or your overwall service. And you have to write to handle a paralle set of code to also handle SIGTERM.
> , but ignoring to handle SIGTERM if there is a use for it is just ignorant.
It seems having 2 sets of shutdown/failure code (one abrupt power failure, "one gentle" is the more ignorant path.
The process that gets sent the SIGKILL can't handle the moment after it gets the SIGKILL. But, the rest of the system could handle the "killing" clenaup (other process and management of resource). That could be another process (a supervisor). Another process an another machine. Or it could be the same process when it gets restarted. First thing it always does is deal with its remnants and after-effects of when it was killed last.
> SIGTERM just lets you do a bit more of the operations you've guaranteed are always safe.
What does that mean? Can you expand a bit. I don't quite understnad "guaranteed are always safe". If the process that gets sent the SIGTERM catches it and does some something "clever". That clever part is never guaranteed. Because the power can cut out, SIGKILL and come unexpectedly (supervisor decides that SIGTERM handler took too long and send), etc.
The application deals with some kind of stream of data. You engineer it so that processing can be interrupted at any point, and on next startup, you're fine--your data is consistent.
Now, once you have that system, I would think that handling SIGTERM is not so hard--you finish processing what's in flight already, while not trying to handle anything new.
From your earlier work, you know that this can't corrupt data. Even if SIGKILL follows while you're trying to handle SIGTERM, you're no worse off than if you just had received SIGKILL in the first place.
The downside seems to be that you now have to show that your system will reliably terminate if it receives SIGTERM.
>The downside seems to be that you now have to show that your system will reliably terminate if it receives SIGTERM.
That's not a problem unless it's followed by a SIGKILL (like from an init system).
What is the point of SIGTERM then. Why bother messing with that code that runs in an exiting/aborting program if SIGKILL will handle. Heck, SIGTERM if not handled will just be SIGKILL.
That is the main point of this -- why write 2 sets of code? Unless you can guarantee your program will always get SIGTERM then SIGKILL. You have to make sure it works correctly with SIGKILL. If you do, then why bother putting in the lines of code to handle SIGTERM?
So no you design for being robust and to handle SIGTERM if applicable. You can never know when you'll get killed, that's why your design have to be solid all the way through.
Supporting SIGTERM and certain other signals is just deciding how big a crater you want to make when the program goes down. That said, it's good practice to use things (write-ahead logging, idempotent operations, etc.) that at least try to mitigate the amount of debris.
Yap that are some of the things I used. Append only mode for files. Code that handle roll-back to last consistent state that runs on every startup. External monitors and watchdog that know how to monitor and to extra cleanup is necessary.
EDIT: In fact they handle a much stronger requirement that a power failure doesn't corrupt data, but the point stands.
Or, if you're on a server, you can use another database. Or use atomic file writes, if you have fancy needs. If you have really exotic needs then sure, do your own thing -- but most of the time, killability is easy and beneficial.
What? I did a pretty simple proof of the safety of Write-ahead Logging as an undergraduate project -- given reasonable assumptions about disks (or implementation safeguards such as pervasive checksumming).
Most developers have superpowers: It's called "existing research". I just wish they'd use them more.
Nope. I did create a lot of complicated programs that work that way. Tested, shiped and got paid for them nicely.
The recovery process after a SIGKILL may require manual intervention, or may leave some data in an inconsistent state. Not every software in the world has the amount of testing for hard kills as an ACID compliant RDBMS - I'd say all the contrary.
So, I don't want to trigger bugs and lose time in complicate recoveries if I am absolutely not forced to do it. Let processes cleanup after themselves, and hit them hard only when and if needed.
I have implemented practical running systems that work like that so it is practical not just academic papers.
> But the real world doesn't always work this way.
Sure, sometimes you don't have choice. You get a 3rd party library/service/db that will shit the bed if something pulls the plug on the machine. You can try to identify those parts and eliminate them or mitigate it as much as possible. But sometimes you can't. Then try to characterise their behavior. Maybe have some external cleanup logic if you care to ensure data doesn't get corrupted. And live with it.
> Not every software in the world has the amount of testing for hard kills as an ACID compliant RDBMS - I'd say all the contrary.
Well but it is something to strive for. If you write the software it benefits to understand what happens to the data. Knowing a bit about file syncing, what kernel does with dirty pages, how/if/when to use O_DIRECT, using append-only modes, a write-ahead log. Think about using checksums in your persistent data, take advantage of atomic file system operations (renames), etc.
Well I certainly saw the benefit and the return from running a better, well tested, well designed service that operates that way.
> Let processes cleanup after themselves,
Think about that statement for a bit. How do you guarnatee processes "cleanup up after themselves". Say a junior tech trips over the power cord of the server, how does the process "clean up after itself"?
But when I'm on my workstation, I need to kill other software as well. I have an UPS on my workstation, and the UPS is very close to the workstation, so the tripping is unlinkely :-)
And when hard crashes happen, sometimes I find some software in an incosinstent state, and I need to cleanup something manually.
Or maybe a remote endpoint for any software would like to be notified of me shutting down, something that's usually done with a clean exit, not with a SIGKILL. Or temporary files would be cleaned up immediately, rather than on a cron-based janitoring job.
* Send SIGUSR1 if TERM doesn't work
* Send SIGCONT if TERM doesn't work
* Attempt to reap the process
* Attempt to ptrace the process
* Give a hint as to why it can't be killed (like blocked I/O, a zombie, etc)
* Kill the process group (dangerous!)
I do unequivocally like the blocking behaviour though, I might implement that for myself.
Edit: so just something like this:
function killwait() {
kill $1
while ps -p $1 &>/dev/null; do
sleep 1
done
}> This tool will hide or allow you to ignore important issues with your system.
(I'm ignoring the fact that a "kill -9" actually still flushes buffers and such, but that's tangential to the point I'm trying to make.)
EDIT: In summary: I'd rather find out sooner rather than later :).
Environments differ though, the software I run will do worse on power failures/SIGKILL than on SIGTERM. I would prefer it to be more resillient, embrace a crash-only design, etc. But my world suffers from practicalities, so I deal with it by being careful with SIGKILL.
Just put everything in one script - these scripts are shorter than an init boilerplate section - and use an option to switch. Do process names by default, then perhaps use a -p arg to switch the script into 'pid' mode. Bash's 'getopts' is pretty easy to implement.
As it stands, you currently have to waste mental power to figure out whether you want the 'n' or 'p' version anyway, so may as well just make it an arg.
ksm `pidof cupsd`
Solving this problem requires kernel support because the kernel is generally the only fault domain in the system that cannot crash while the system is still running.
Edit: forgot the link the first time. [0] http://illumos.org/man/4/process
http://lists.dragonflybsd.org/pipermail/commits/2014-Novembe...
So now FreeBSD and DragonflyBSD have plans to utilize this exactly like Illumos and build lightweight process supervision without the need for something as disruptive as systemd.
> [..]
> Bash
Maybe I'm nitpicking here, but why Bash _should_ be available?
Linux