Linux gets frozen, what do you do?
jovicailic.org
jovicailic.org
It wasn't even mandatory to restart the system: the trick was to use first MagicSysRQ and then issue a "vgareset" (IIRC I couldn't even see what I was typing, but the command was taken into account) and then, miracle: I could unlock frozen X Window sessions (more specifically: kill X, "reset" the GPU and then restart a new X Window session).
Note that very often X is fine: it's just X which is frozen. Heck, if you have another machine on your LAN and allow SSH in, you can SSH and kill X / vgareset without needing MagicSysRQ.
But since quite a few years X is so stable that such hackery ain't needed anymore. Moreover if I recall correctly "vgareset" did only exist for 32 bits system (at least at one point). Nowadays my Linux workstation regularly reaches 6 months of uptime (there are only very rarely know remote root exploits mandating a kernel upgrade) so I've kinda "forgot" how MagicSysRQ works ^ ^
You can force a kernel panic by 'echo c > /proc/sysrq' if you really want.
http://en.wikipedia.org/wiki/Magic_SysRq_key
Mind you I don't think I ever used "Stop A" on a Sun for anything constructive...
Ahh, the good old days... I once implemented a nice little boot device selecting utility in for for some of my Suns. All written in Fourth, and executed by the firmware on every boot.
>b -s
and then, at the sh prompt, as the single (super) user, edit password # ed /etc/passwd
to add in a second root account for yourself. That could be pretty helpful.Back at the dot-com I worked at, we used to have a SPARCstation 5 that we used for SPARC-II arch builds. It was a nice machine, but it had a dead NVRAM battery so it would lose its configuration every time we lost power in the building (which was a lot, because rolling blackouts used to be a thing in California).
Anyway, an interesting quirk about the SPARCstation machines is that the MAC address for the NIC is stored in NVRAM, and when you lose the NVRAM settings it defaults to ff:ff:ff:ff:ff:ff AKA the broadcast address. So by default, on boot, this machine would start sending out DHCP requests and other network traffic with a source address of the broadcast address. The switches we were using did not like that, and would start flooding all of their ports with traffic. The only way we found to fix it was to reboot the switches.
So, there is at least one constructive use for "Stop A": you can use it to configure a MAC address on your SS5 so that it doesn't inadvertently bring down the whole network in a massive broadcast storm.
You mean "Linux" is fine?
In any case, in ~8 years of using Linux, I don't think I've had a freeze that was the kernel and not just X.
But agreed on 'just reboot': you should be using a resilient filesystem anyways, and MagicSysRq only works when it's enabked beforehand, and only for some kinds of lockups, and if the data is in the frozen application, syncing the disks isn't going to help.
Yup. Changed my GPU to nvidia, no more core lockups.
> I don't think I've had a freeze that was the kernel and not just X.
Device drivers sometimes have bugs, especially when the device is not working properly (I had system freezes when there was a misbehaving capacitor on the graphics card).
Wouldn't that make it an unsafe database server?
Edit:
Besides, hard drives have write caches, and they can report a successful write operation to the OS when the data is still in its cache physically.
Hard drivers are supposed to flush their caches before they report the end of a flush operation to the OS (some flush into flash, but they flush). If your does not, it's defective. Go ahead and make use of the warranty.
No properly configured and working ACID compliant RDBMS should lose any data when the server is reset or stopped. If it does, then it is either a problem with the hardware, OS, configuration or the RDBMS itself. The application must also be able to handle the DB disappearing, though. Sadly this is often not the case.
I remember some sort of magic invocation like that years ago for our medical "device" product we had to manage hundreds of remote node instances of. Something along those lines. I don't remember why the developer (Gabe) came up with sync three times being magic number. Maybe three times was just paranoia. =)
Somehow this advice got mutated so that you'd keep holding the keys until you heard two boot chimes (thus resetting the stuff twice). And then it started to grow. Three was common. Some people would advise more. I'm pretty sure that doing it more than once never helped anything, but there we are.
(The cmd-opt-P-R sequence still works on modern Macs and I actually used it to resurrect a machine that wouldn't start up just a month ago, but it's far less frequently needed now.)
Or the "Repair permissions" thing in OS X. You do it several times as well.
It's like whenever there's this one-step fix thing that a system utility does, the Common Man will interpret it as needing to repeat 3+ times in order for it to be effective.
"sync; sync; sync; halt"
Presumably so you had thinking time before automatically restarting a possibly sick system.
When your Sun is particularly hosed we used sync;sync;sync; halt to reset and (hopefully) not lose any data (sync forces OS write buffers to purge)
>Run it and be amazed how much your disks/raid/OS lie. ("lie" = an fsync doesn't work)
>It seems everything from PATA consumer disks to high-end server-class SCSI disks lie like crazy. Yes, that includes SATA there in the middle. I'll discuss fixing your storage components in a second.
anyway, someone already mentioned a few mnemonics, I learned one, a quick googling lead me here:
http://fosswire.com/post/2007/09/fix-a-frozen-system-with-th...
For example, on a Linux laptop with the hard-drive encrypted with dm-crypt, I simply lost access to my drive due to repeated hard reboots. I don't know if I could have recovered my data from it or not, but after repeated attempts of googling for the error message and following advice I simply gave up and later reinstalled everything from scratch (it's a good thing I constantly make backups ;-)).
On Linux REISUB has been my friend.
1. Tap the power button (one single brief tap)
2. Wait a minute
3. If the system hasn't already shut down or isn't obviously in the process of shutting down or you just get impatient: hold the power button down
4. Now that the system is off: tap the power button.
Modern systems (whether server, desktop tower or laptop) have a "soft" power button, that when tapped briefly sends a signal to the OS. Most flavors of linux are configured so that receiving that signal initiates a proper shutdown, just like "shutdown", "poweroff" or a GUI shutdown. All of this has been true for quite a few years now.
If the system is so locked up that the power button initiated shutdown doesn't work, you might as well just pull the power, which holding the button down for a second or two will do. Even if this happens, you're probably using a journaling filesystem that comes back up cleanly.
None of that will help with unsaved data in an application, because you almost certainly have to tell the application to save, and if X has locked up there is no way to do that. A clean shutdown might signal the app with a terminate signal and allow for data to be saved before doing a hard kill.
echo 1 > /proc/sys/kernel/sysrq
echo b > /proc/sysrq-triggerYes, it's exactly like pulling and replugging, which was exactly what I needed!
--edit-- is sync even a bash builtin? Looks like it's /bin/sync on my systems. / had been lost (it was on a usb stick on an internal header on an unreliable usb3 card, I later found out.)
--edit 2-- if you meant also echoing s to the sysrq-trigger, it seemed to kill the session
How could that happen and leave your system in a state where it is still accessible and can be fixed by a reboot?
If you think you might face that kind of trouble you should keep a copy of Busybox and whatever other tools you might need on a RAM drive. You'd have an opportunity to figure out if rebooting would lead to a usable system.
There was a problem with the card or the driver as every so often something would go wrong and everything would stop responding again. The shell I left open revealed almost nothing as the root drive was gone, but it could be used to reboot the machine (thanks to the trick above), which would then be good again for another few days.
I now have / on a proper SSD...
Use acpi to set your power button to switch you to console. Just edit your /etc/acpi/power.sh and change it to...
#/bin/sh
/bin/chvt 1
The acpi system is separate from the X/keyboard interface and still works. So this way you can switch to the console even if your keyboard is locked.Note you can use other keys as well. On my laptop I use my "thinkvantage" key for this. Any key that has an associated acpi event will work.
$ cd /etc/acpi
$ cat events/thinkvantage-button
# Thinkpad 'ThinkVantage' button
event=button/prog1
action=/etc/acpi/thinkvantage.sh
$ cat thinkvantage.sh
#!/bin/sh
/bin/chvt 1
Hope this helps.That, and because X consumes those keys, making them useless when X goes bad (what is about all the times that Linux freezes and it's not hardware fault).
Section "ServerFlags"
Option "DontZap" "false"
EndSection(...from painful experience...)
It's the "Secure Access Key" (SAK): You press that key and it kills all programs hooked to the TTY (incl. X, in your case) and displays a proper login prompt so that you can know what you're about to login to was run by the system and not a clever malware trying to steal your password.
This is a laughably awful shortcut for anything. Why make the supposedly safe stuff so out of reach? I would have to use a separate computer to look up how to type this command. Even if I knew the command off the top of my head, what if my keyboard lacks a PrtSc (SysRq) button?
When folks are pointing to Windows for design cues, the open source community ought to reconsider some aspects of how it has implemented certain features.
Linux, by the way, can also respond to Ctrl-alt-del, and depending on the desktop it might perform functions similar to Windows.
I had to use this just today when ddd stole all other mouse and keyboard input.
You can enable them in the keyboard options of your DE, or directly with setxkbmap on startup (sorry can't look up the exact options right now) if you're not using one.
of course, fixing why your X hung would be wise too, since for any normal distro and mainstream hardware, that's just not going to happen. if you really want to be fubar, have that f@cking POS systemd die on you... (yes, on-topic, since only sysrq saves the day.)
In more than a decade of running Linux on my desktop (and at work), I genuinely can't think of a single instance when I've not been able to pull a virtual console from a frozen desktop (albeit it often performs laggy).
[0] I'm actually not 100% sure about that. Thank God I don't have to worry about it anymore.
Secondly, I was making a general comment about peoples desktops rather than talking specifically about your example (given the lack of details you posted, it would be insane of me to assume I could diagnose your fault with any precision). My point was that generally when people think their computer has locked up / X has crashed, it's actually one of the items I mentioned earlier that's at fault.
The snappy reply was appreciated though </sarcasm>. But given just how unusual your circumstances are (assuming what you said is true) and how much you seem to hate it when others discuss these topics with you; it might be an idea if you clarify your position a little better the next time such a topic arises. Like maybe saying "my crashes aren't typical because I'm a kernel developer, but.....". This way people don't accidentally post something that hits one of those raw nerves you have and it saves us all from a lot of unnecessary condescension.
Each one of these commands (r, e, i, s, u) takes or can take a few seconds to complete successfully, so let them do their thing.
In particular, if you have HD-activity LEDs, watch and wait until they stop blinking - after the 's' (sync) key.
unused The server encountered an internal error or misconfiguration and was unable to complete your request. Please contact the server administrator, (etc...)
Did their Linux get frozen?
[edit] it works again
This is especially a problem with Docker, I've found, because when I've got a container thrashing the system like this, I can't access the host or any other containers.
Is this because I'm running single-core instances and the problem goes away once I have more than one thread available? Or should I be restricting workloads to a block device not used for caching? Does anyone else experience this?
This throws away one of the neat advantages of containers compared to VMs, though, which is that you can overcommit (or "thin-provision", if you want to be charitable) the host's resource allocation, since it's extremely unlikely that all your services will balloon in resource-consumption simultaneously.
A less guaranteed, but more economical strategy, is to just set quotas on each container such that if it individually starts using more than, say, 80% of the host's resources for an extended period, it'll be terminated. This doesn't save you from a bad interaction between containers that makes everything explode in parallel (e.g. all your containers attempting to reconnect to a stuck service with no backoff) but it should save you in the majority of cases where each host has a heterogeneous container-load, and horizontal scale happens between hosts rather than internal to them.
(The dev-side solution to this problem, though, is to just set up your software architecture such that any host that gets into such a state can hard-reboot, without you losing anything of consequence. http://12factor.net -type containers excel at this, and it looks to be the strategy https://coreos.com embraces as well.)
I don't know if it's that something Docker can fix. It might be something kernel devs can fix, though.
Also, you don't need to go all the way through it. In my above example, MagicSysReq+RE or MagicSysReq+REI is enough usually. E and I are for sending SIGTERM and SIGKILL, so all the processes are now dead, and we have plenty of RAM. Sometimes it does take a bit of waiting: the system is disk thrashing. That said, you can then just resume, though you'll likely need to switch to a VT and restart X and other things. (But you get to keep your uptime.)
There's also F, which invokes the OOM killer, though I often forget about it.
That's going to be faster than the time spent to search google for the magic incantation.
... to be honest, I can't think of another scenario. I have an Advantage for every place I spend extended computer time in -- it's too big and clunky to carry around...
Any ideas?
My 8yo has a mnemonic for it something about Elephants and Umbrellas but I've always done reissub (now with extra sync'ing power). For a time the process for him was login, Alt+F2, konsole, Ctrl+R, mine, Enter, AltGr+SysRq+R, E, I, S, U, B; then repeat the first part and you're finally ready to play Minecraft!
It should print something like:
"SysRq : HELP : loglevel(0-9) reboot(b) [...]"
sysctl -w kernel.sysrq = 1
Stuff kernel.sysrq = 1 in /etc/sysctl to make it permanent.I do not really understand why it isn't enabled by default though...
I think I need an extra hand to do this
Maybe it's easier in saner keyboards.
This is the layout that I have (and the only keyboard layout I'll ever buy, because every other throws the pipe key and backslash in a random spot along with randomly sizing the enter key)
http://commons.wikimedia.org/wiki/File:ANSI_Keyboard_Layout_...
I'll need three hands.
If a Windows box froze, and you had a (somewhat slim) chance of gracefully shutting it down, would you not use it?
Would you call pressing F8 to access magical boot options in Windows, a reason non-technical people wouldn't use Windows?