When should I not kill -9 a process?
unix.stackexchange.com
unix.stackexchange.com
If you're a developer, before you kill -9 a program send SIGTERM (ie kill without args or kill -15). If the program does not respond, run gdb -p <pid> and then "thread apply all bt" before killing it. At the very least, you should get a good idea of why it was not responding to other signals.
Saves a lot of hassle with manually attaching GDB.
* catching of NULL pointer references (ie. transforming SEGV into catchable userspace exception) * testing whether some in memory data have been changed/accessed (ie. mprotect(), wait for SEGV, set flag, un-mprotect(), return) * paging in userspace (eg. across network)
All that seems obscure (and is mainly useful for virtual machines and such), but it is not so rare to find process that actually does this (usually because of some library, or because it is some kind of virtual machine)
I think it was Oracle.
If I remember correctly we had to get the vendor involved
This happened after an unclean shutdown (power failure). The disk were RAID, but the cache battery was dead so there was some corrupt data written to the disk on block level.
THis could happen too on a kill -9.
At the very least, you need to be prepared for a possibly very long replay of logs. Also, a huge number of people run databases in a configuration that doesn't make those guarantees. For example, many people will run mysql (esp. less performant slaves) with innodb_flush_log_at_trx_commit = 0 for performance with the understanding that a failure might require manual fixes.
Losing a transaction that was committed upstream can lead to a bunch of manual fixes in replicated database clusters. What happens if rows inserted in that lost transaction are eventually updated? Replication breaks, you get paged, and now you're faced with either manually fixing the DB consistency error or by rebuilding the whole slave. Woe be unto you if this host is a replication hub with a bunch of slaves hanging off it.
Edit: Especially when you have multiple systems that should be consistent. Deleting a user from the DB, but being killed before it can be removed from our auth system, for instance.
Postgres has special code to account for hard kills: it has a SYSV shared memory segment, to which every child attaches. If the parent dies, the children don't have a good way to know, so they might keep running. If you try to start a new postmaster and it sees that there are still processes attached to the shared memory segment (shm_nattch), it will fail to start.
Were it not for that code, it would allow you to start a new parent process, leading to chaos as two sets of backends were accessing the same files and shared memory without knowing about eachother.
So, software should be written to assume it might be killed if you want to have a robust system.
By the way, using a small fixed amount of memory is no defense, or at least not in all kernel versions. The "badness" heuristic function used to find the victim could end up counting the same byte of memory many times:
kill -15 typically leaves processes in a properly shut down state, which in terms of databases alone means that they will start up without a recovery process (which can be a 20-30 minute operation sometimes). That alone makes waiting a few minutes for a running process to respond to a kill -15 worthwhile.
For admins, I wouldn't fault them much for kill -9, I would fault the software more if it lead to anything more than an inconvenience. But sure, it's wise to use -15 or whatever as long as it works.
If your process relies on not being kill -9'd, then you might as well quit programming and go buy a lottery ticket.
As infrastructures have grown, and I have managed large applications involving tens to hundreds, often over a thousand servers, and I have grown to accept that a power supply can fail and a node can disappear from the network and it's even possible that none of its' components, including its' drives, will ever work again. I've never _really_ experienced such a catastrophic failure, but it's a lot easier to sleep at night if you just assume that.
kill -9 should never be worse than pulling the power plug, which is what netflix's chaos monkey always tries to simulate.
we all have to live on a continuum of how much of that we can survive, but if you always assume abrubt failure, it'll be pretty tough to give you a bad day.
No, it shouldn't, but just pulling the power plug isn't exactly recommended behaviour, either. There aren't many admins out there who will happily yank the power cord out of their desktops when they want to power it down.
Last night I spent several hours getting a server back into gear after a 'pulled power plug' event. A friend's rack was affected by a (seven-hour!) substation power outage, and it wasn't on a UPS, so it didn't shut down cleanly. Eventually we were able to coax the server to boot again (had to remove all USB devices in the process, including internal ones), and the problem was a corrupted MBR. Make a rescue usb stick, boot into that, finally diagnose the problem, cat a new MBR onto it, and tada, fixed. Let's just say that I don't find the argument "shouldn't be worse than just pulling the plug" to be particularly comforting at the moment :)
I've had abrupt shutdowns happen on laptops (one had a particularly loose battery...), desktops, servers and although have encountered corrupt files and filesystems, never had any of them corrupt the MBR.
I know that grammar comments are probably not welcome here on HN, but I think that since you seem to have an interest in it, I'd point this out:
"Its" only has an apostrophe when it's a contraction of "it is" (or "it has"), the possessive is always "its" (without an apostrophe).
</offtopic>
It's actually easy to misdesign a system such that it is safe against power failures but not kill -9. See my other comment:
http://stackoverflow.com/questions/2405172/resources-about-c...
I have heard the best practice advice for many years and I think that the 'you should send some friendly signal first' is not universally what works out best. For instance, if your Chrome browser is getting out of hand and the system is permanently doing some 96% wait for some reason, a gentle killing of Chrome will take ages and, when it restarts, you might get some but not all of your tabs back. With a killall -s 9 you can be back to work quickly with all your tabs (and underlying swappiness problem hopefully resolved).
Sadly, what he failed to check was disk IO stats, this MySQL setup had heavy innodb table usage and settings that where deliberately set for more performance then reliability (large buffers, delayed commits etc.). What was going on was normal, MySQL was flushing everything to disk and to logs and was most likely going to stop without a problem.
He didn't look at the facts at hand, the disk IO was still going, MySQL was mostly writing to the log files, users where not being let in so the db was doing an orderly shutdown. Instead with the adrenalin pumping he felt he had “waited a long time” and issued the kill -9 and corrupted the InnoDB logs and tables beyond all recognition.
I landed at the airport to five frantic voicemails because this db was the core of a bunch of high profile sites and he was up to his ears in phone calls from the client. I had to spend the first 9 hours of my vacation with my kids playing in the background while I sweated it out on a laptop over a crappy connection that kept dropping me.
Yes I know, MySQL should have been able to handle the "power out" but this event was made worse because he started the shutdown, we had a deliberately fragile implementation, he didn't check the slaves so we didn't have a clean fall back and meanwhile he "waited a long time" but never checked the process to see what it was doing.
I use kill -9 (-KILL) all the time, but I do it where I know it's needed. Most of the time kill just works and if it doesn't that should give you pause to think carefully about what you'll do next. Slow is Fast and Fast is Slow, if you quickly do something radical like kill -9 or init 6 or 10 second power button crash then you may be spending the rest of your day cleaning up. Slow down a bit, look, listen and gather facts about the situation then make an informed decision. At least if you do all that and the rest of your day is still ruined you won't have that nagging feeling you shot yourself in the foot and you can talk intelligently to your client or boss about steps you took to avoid the situation.
My failure I suppose was that I hadn't explained to him that it was typical for the db to take upwards of 5 to 8 minutes to cleanly shutdown. Which gets to a second topic, documentation for production systems is essential, when the fire is on too many mistakes can be made because of "knowledge gaps" between team members. Needless to say after this incident I wrote extensive documentation for the team so the next time I was "on vacation" I could actually be "on vacation". :-)
Postgres is designed to be resilient to "kill -9" as well as hard power offs. Even if you use durability-sacrificing features like asynchronous commit[1] or unlogged tables[2], the risks are very well-defined and contained to recent transactions and data in that unlogged table, respectively.
But even for postgres, you have to be a bit careful. For instance, many disk drives lie about completing the writes and really have them in a volatile cache, so a hard power off can still cause corruption. You need to disable the write cache on the disks using hdparm (or similar) to be safe. And "kill -9" is quite annoying, because the child processes don't have a good way to know that the parent is exiting, so then you will be unable to start the new parent process until the children have all exited as well.
EDIT: There's still no excuse for MySQL completely corrupting the system on a kill -9. That's just a misdesign -- consider that the out-of-memory (OOM) killer on linux sends a -9.
[1] http://www.postgresql.org/docs/9.4/static/runtime-config-wal...
[2] http://www.postgresql.org/docs/9.4/static/sql-createtable.ht...
If the database truly corrupted only due to improper shutdown, it was because the double write buffers were disabled on a non-atomic FS in which case you're intentionally risking DB corruption. More likely, there was bit-rot which was undetected during the normal running state of MySQL, and recovery couldn't get around it.
On the anecdote side, I've yet to have MySQL corrupt a database from a kill -9, and as I do a lot of failover development and testing, I'm issuing kill -9 to running databases frequently.
I guess the moral here is tread lightly and take the time to be informed before you just smack something with a club.
Surely just sending SIGKILL will allow the writes already issued to complete, regardless of the FS type/options? Nobody kicked the plug out of the wall.
InnoDB writes to the doublewrite buffer (an internal tablespace on disk), then the record in place. If the two do not match on recovery, the record is considered not written, and rebuilt from the transaction log.
If the doublewrite buffer is disabled, then InnoDB has no way of telling if the write to the page on disk was started, completed, or partially completed. This causes corruption.
The advice about Slow is Fast in such situations is spot on.
Reminds me of a Unix incident at an automobile component manufacturer's plant. They had a multi-user Unix box supplied by the company I worked for at the time - a large Unix hardware and software vendor. A colleague and I had gone there for some system maintenance. In order to do something, my colleague, before I could stop him, gave the command "init 0" on the root console. ("init 0" shuts the system down without confirmation, for those who don't know, if run as superuser.) Within seconds the phone in the computer room was ringing madly, with calls from different intercoms on the shop floor, inventory store, etc. He had to apologize to many of them before they calmed down ...
Was the extra performance in any way worth it?
Sadly it worked very well except when people used a big hammer on it. We had replication slaves but in this incident the salves were not replicating and somehow the script checking them wasn't alerting. But even so the admin didn't check before bashing the keyboard and thus we were left with manual reconstruction from backups and other sundry sources of information.
So, that was by the client's demand. You got to save 20 machines with it, I'm impressed, and now I understand it better why you did it.
I can go on for hours about those two words, everyone wants to be their own boss until they don't want to be... :-)
The only reason I can think of to shutdown a production SQL database would be to do maintenance. In which case any sane admin should be taking a backup in advance. That should be the first step.
If the database is not a production server, does it matter if data was lost?
If the production database crashed/bugged out and therefore needed to be halted, was it really ready for production in the first place?
I don't mean to offend, but it sounds like your partner is a bit of a doofus.
The core issue was actually a deadlock but he wasn't trained in detecting that and in the past just restarting the db "fixed it". Because of course all the sessions dropped and dead locks were resolved.
This was very much a high visibility production environment and yes we all make mistakes sometimes.. How else can we learn.. Trust me he never did that again! :-)
Remember, all of use were a doofus at one time and even worse many of us are every day.
It matters very little to me how people screw up, it matters how they recover, learn and avoid doing it again in the future.
My favorite saying "Originality in mistakes" if you're going to screw up make it big, creative and interesting. And please, don't repeat it.
Sure it will mess up some things, but when management is pushing to limit the downtime of, for instance, a golden-image provisioned Linux machine, I'd kill it off no problem.
Now when we are talking a hardware box running some form of Oracle/mySQL - no, don't use -9 indeed.
In my case, processes I have to `kill -9` are
- programs that only read data files (or write to dispensable files)
- programs that I know don't react to SIGTERM (there is no cleanup logic, but still something makes them swallow SIGTERM)
- often then are simple tools (e.g. ls) that become wedged in a system call or in kernel code (when trying to access a bad NFS share)
- in the other cases they are in-house programs that are either badly written, or too complex (the worst offender is CERN's ROOT, if it becomes wedged you have to `kill` and `kill -9` several processes it spawns), or where we don't care enough to fix them
Interestingly, there seem to be some cases where even `kill -9` doesn't help. What I do then is to freeze the process with Ctrl+Z (Ctrl+C doesnt work of course), and then `killall -9 $(jobs -p); fg`.
Actually, I have one program I routinely call with `program; killall -9 $(jobs -p); fg` and end it with Ctrl+Z. Sad, but true.
(Of course, if your process is a database or a GUI tool or something, then all the standard wisdom against `kill -9` applies.)
kill -HUP <PID>
kill -KILL <PID>Unix one-liner to kill a hanging Firefox process:
http://jugad2.blogspot.in/2008/09/unix-one-liner-to-kill-han...
It had an interesting thread of comments in which both others and I participated, and at least I learnt some things.
At least that what I do.
I guess a process is given more space to "clean up after itself" with a normal `kill`; where a `kill -9` forces it to die.
Anyway; I don't know the exact answer -- will come back later to read a wiser person's answer. :)
Never.
WRT shorter: Magic numbers don't just suck in programming.
7th Edition AT&T Unix did not allow signal names (http://plan9.bell-labs.com/7thEdMan/v7vol1.pdf, search for "extreme prejudice" -- I still remember many of the little gags in the early manpages).
That was a mainstream release in the mid-1980s. Even the basic utilities like kill(1) were incompatible back then, so if you worked on both BSD and AT&T systems, it was easier to use the compatible subset.
*
Regarding magic numbers: in general, yes, to be avoided. But my usual use of kill -9 is in exasperation, from the command line, and clarity for others is not a priority. I admit, in a script, kill -HUP is to be preferred to kill -1. But even in a script, I'd say kill -9. This usage thing seems to be complex.
~ The most interesting man in the world.