Why not kill -9 a process?
unix.stackexchange.com
unix.stackexchange.com
There is a lot of middle ground between failing "permanently and dramatically" and shutting down cleanly.
What about a program that can recover after an abrupt termination, but only after a time-consuming recovery from a journal? What about a program that can recover after an abrupt termination, but only after you manually remove a lock file? These are cases where it's not "just fine" to use SIGKILL, but not so bad as to warrant removing the program.
Generally your life will be easier if you don't take the sledgehammer approach to killing processes, and at least try a non-SIGKILL first.
If a process doesn't respond reasonably to SIGTERM, then you should consider removing it from your filesystem.
(I've supported clustered Oracle on AWS handing massive numbers of micropayment transactions and seen weird shit in prod where it "kind-of failed" according to our ops dba. Classes of survivable bugs vary from apply a vendor hotfix to edge cases not worth downtime.)
What's reasonably? MySQL/InnoDB will start consolidating the journal and buffer pool in preparation for a shutdown. On machines with a large amount of RAM allocated to the buffer pool that will take quite some time. Is that still reasonable?
InnoDB is designed to recover gracefully from a SIGKILL, too. The journal task is "saved" for later. That's why they tell you in DBA school, "don't use MyISAM tables. They're not safe." Because if they received a SIGKILL or power outage at the wrong moment, writes could have been lost. Amiright?
What you don't expect is for a process that receives SIGTERM to fail hard, and require an expensive journal recovery or suffer some unrecoverable data loss as a result.
I didn't know that.
https://www.usenix.org/legacy/events/hotos03/tech/full_paper...
It should be easy to rationalize that SIGKILL recovery is slower than SIGTERM shutdown for such a database. If your in-memory cache is empty, it won't be providing any speedups, right? Thus you'll need to go back to disk for everything.
He: You mean SIGKILL isn't the signal for graceful termination? And it isn't sent by default when invoking `kill <PID>`?
Me: Nope. Why do you think that?
He: Well, the command is called `kill`.
Me: ...good point.
That's an overreaction. I've seen programs refuse to respond to SIGKILL for reasons out of their control, like getting stuck in an NFS transaction or other "unusual" filesystem and the kernel being unable to process the SIGKILL. "Listing a directory" is not exactly a crazy thing to do.
And no wiggling out by saying "well, that's the kernel, not the process", the process never gets a chance to "handle" SIGKILL, so it can't really screw it up, either. Arguably, all such failures are the kernel; the kernel should not expose any sequence of calls that causes SIGKILL to fail, and over time they tend to be fixed (I haven't seen this on my modern Linux machine in a long time, even playing with some funny stuff), but it has happened and will probably continue to happen as new stuff comes out.
Technically, you are correct- the process doesn't even know it's been SIGKILL'd. However, there are other things it could do to gracefully handle the scenario upon next start. Is there an already-existing lock file? Prompt for its removal. An already existing PID file? Again, ask what to do. Include tools to fix records that may have been left in an inconsistent state. So on and so forth.
That said, I've never had to do a SIGKILL. However, I still expect my programs to sanely recover from power outages, clumsy interns, RAID failures, and other "acts of god" that may suddenly cause a program to end before it has a chance to clean up after itself. It's part of making robust programs.
I currently have a vim process that is running in the background and not responding to SIGTERM (so I'll probably SIGKILL it). Does this mean I need to find a new editor?
*BSD with ZFS.
also -9 can leave ipc hanging so some ipcs might be needed.
It would be far too easy to kill any program mid-write().
Unless you're talking about a database not intended to be used in production, that's really not fair.
What if the power comes down? The UPS blows up? There's an earthquake and the ceiling comes down? Is it OK for a DBMS to be corrupt its data then, too?
Yes it is. The DB should not unrecoverably corrupt itself in case of hard power loss, and a kill -9 is significanly less traumatic than a hard power loss (which includes PSU melting/explosions or UPS going down).
I'm being a bit pedantic, but trading time at shutdown vs. time at start-up + potential data loss (if the journal was in the middle of being written while the -9 was sent, you probably lose the last journal entry).
I disagree. If a journal write was in progress on SIGTERM, you either fail or complete the write, returning the result to the client. In either event, you should end up with a clean journal on SIGTERM, assuming your implementation was sane (i.e. how I'd write it). If the client chooses to ignore the failure, that's not on the server and that's the whole point.
A normal shutdown should never require data loss. Somehow you're conflating SIGTERM with SIGKILL and I'm not understanding why.
And the whole point of a journal is that it doesn't need to be clean. If your database requires the journal to be clean on startup as a condition for its correctness, then the journal is useless, it could just as well just require the data itself to be consistent. The only reason why there is a journal is so that correctness is not affected when writes are interrupted as any point.
Also, for that matter, a SIGTERM might be able to guarantee that the journal is clean, but it would be highly broken if it tried to guarantee that the client gets to know about the result, as that could take an arbitrary amount of time that might be dependent on the behaviour of remote systems, which would be terrible shutdown behaviour indeed.
Really, the only reason why a database should even try and catch SIGTERM is when checkpointing on shutdown is a lot cheaper than log recovery on startup - other than that, it only makes the code more complicated without providing any benefits.
hangs head in shame
It wouldn't be unexpected for a SIGKILLed process to fail to flush a cache, or to exit with persistent state ...inconsistent. That can be permanent (changes lost) and dramatic.
You can wrap operations in transactional overhead to greatly reduce the chance of data loss, but you can't eliminate it entirely without assistance from peers.
kill -9 is the equivalent of unplugging the machine in the middle of operations. It might be fine, and with careful design it might be fine almost every time.
And why "might be fine almost every time" instead of "will always be fine"?
There are several kinds of bad state that can be left over:
(1) space leaks (disk, memory (sysv shared memory))
(2) inconsistent data (data / config files)
(3) locks/semaphores (used to protect (2) from happening in normal operation)
(We're talking about those that are under the control of the program itself, not the control of another program or the kernel (the latter can and generally will be cleaned up automatically).)Programs can be written to clean those up when they are restarted, which is the reason for the saying that if some program does not do that, then it's not worth being left installed.
But it takes a certain amount of work writing code to recover from an abandoned lock and the possibly inconsistent state that has been left over, so much so that even some wide-spread system libraries don't do it. Of course you're free to remove those system libraries from your system, but you may have to write a replacement.
Also, in the case of such libraries, it may happen that you lock out other programs (that also happen to use the same libraries) than the one you SIGKILL'ed, and you may not be aware of the source for the issue.
In particular I'm thinking of ALSA. If you SIGKILL a program while it uses [a "hw:" device in] ALSA, it leaves behind a semaphore that prevents other programs from accessing the same device indefinitely (it seems there's no recovery code in the ALSA libraries). This robbed me of quite some time to figure out; when I did, I wrote a utility to clean up the semaphores explicitely[1]. Every now and then it happens that I need to reach out for it.
I had sort of forgotten the importance of the signal argument. Such an incredibly powerful command that is nothing more than a representation of a very simple decision or architecture [2]:
Some of the more commonly used signals:
1 HUP (hang up)
2 INT (interrupt)
3 QUIT (quit)
6 ABRT (abort)
9 KILL (non-catchable, non-ignorable kill)
14 ALRM (alarm clock)
15 TERM (software termination signal)
[1] - http://www.youtube.com/watch?v=Fow7iUaKrq4[2] - man kill
kill -9 -1 to send a KILL to every process.
% kill -L
kill: unknown signal: SIGV
kill: type kill -l for a list of signals
% which kill
kill: shell built-in command
% /bin/kill -L
1 HUP 2 INT 3 QUIT 4 ILL 5 TRAP 6 ABRT 7 BUS
8 FPE 9 KILL 10 USR1 11 SEGV 12 USR2 13 PIPE 14 ALRM
15 TERM 16 STKFLT 17 CHLD 18 CONT 19 STOP 20 TSTP 21 TTIN
22 TTOU 23 URG 24 XCPU 25 XFSZ 26 VTALRM 27 PROF 28 WINCH
29 POLL 30 PWR 31 SYSYou can disable kill (or any built-in)
Under zsh:
disable kill
Under bash: enable -n killEDIT- at least it looks like the calls are the same.
SIGKILL shouldn't be your first resort but sometimes is necessary as a last resort. If the sky falls and you can't deal with it, there's something wrong in a lot of places which have nothing to do with signals.
Usually it's just an obviously good idea to send a milder signal first, because it's less likely to leave orphaned child processes and `kill 345` is just plain easier to type than `kill -9 345`. Also, trying a TERM signal first gives you some feedback on how fubar the process really is.
For example, postgresql forks a process for every connection. What you may not know is, if you kill one of these processes, it needs to clean up it's use of the shared memory pool. If you kill -9 any of postgresql's child processes, the other processes will see that a peer died uncleanly, and postgres will just shut down rather than risk corruption.
Oh, the things we learn the hard way.
So pretty sure this is totally fine and not a problem in the slightest, seeing as some billion devices or so are doing this on a daily basis.
You do need to be a little careful as an app writer. All your file writes should be atomic (either SQLite/CoreData or using the atomic write methods in CoreFoundation and Objective-C) or you must be able to detect and delete partially files and recover in-situ.
You should be doing that regardless, there's nothing special about SIGKILL in this regard. Dirty pages will still get flushed to disk, so SIGKILL is safer than, say, sudden power loss. Which isn't exactly unheard of on battery powered devices, after all.
Background services have weaker lifecycle guarantees, but the system can be asked to automatically restart your service and re-delivery any Intents that were being processed.
It's not like the system is regularly killing apps without any recovery options. Dealing with these lifecycle events is a key part of Android development, so apps are designed to deal with them.
See: http://developer.android.com/reference/android/app/Activity.... http://developer.android.com/reference/android/app/Service.h...
But that call only gives that activity a chance to save state, not the process as a whole. You don't run around stopping all your services and such in onPause(). There is no SIGTERM equivalent where you can go around doing actual cleanup work for the entire process.
As for Services you'll note it's considered killable at any given point, there's largely no guarantees other than it won't be killed in the middle of executing code in onStart/onStop. But during the bulk of its actual work it's totally up for random killing.
And fwiw foreground processes are totally considered killable. They are the last in the queue, yes, but they are still in the queue.
On standard UNIX there's no onPause() or anything similar, so the process cannot react to this in any way.
Also to be clear the onPause and SIGKILL are not tied together. You could get an onPause and it be minutes or hours or even days before you get SIGKILL'd, during which you are completely free to keep running code in the background.
And depending on how your process was started there might not even have been an Activity to be onPause'd in the first place. Consider an app that started doing work in response to a broadcast or content provider query.
"I use kill -9 in much the same way that I throw kitchen implements in the dishwasher: if a kitchen implement is ruined by the dishwasher then I don't want it."
https://gist.github.com/anonymous/32b1e619bc9e7fbe0eaa
(The cleanup might have introduces some subtle bugs, but a quick test showed no issues. Let me know if you find some!)
temporal@legion~/r/mintia> murder 4421 0 50/50℃ 13:56:10 28.01.2014
murder: Killed process with PID 4421 with signal TERM (15).
Works like charm :).Was this submitted to HN for more opinions? So people could see the disagreeing answers? Something else?
And vice versa, remove any program where -9 may cause damage.
I kill'd the script running it, and then kill -9'd it when that didn't work. Two weeks later someone asked about my query that was still running on the database.
And now I'm the one who warns people not to kill -9 scripts without understanding why it's stuck and how to clean it up properly.
If nothing else works, kill -9 -<master pid> to kill the whole process group, otherwise detached processes owned by init could get messy.
By far the easiest question to weed out inexperienced linux users.
This is my favourite part of dd: The "will I receive a status report or will I terminate my long-running copy?" gamble.
Or use sigchld to a parent to get rid of zombie children.
It is better for the parent to catch a TERM and relay that signal to all the children. I find this to be a practical and typical use case...