Use this kernel parameter in your kiosk
cedwards.xyz
cedwards.xyz
Is it perfect? No. But you can do much, much better than “is the kernel up.”
Ie. every 30 seconds, contact an update server server and, if successful, poke the hardware watchdog.
That means every device in your network will either be rebooting, or correctly talking to your server with the latest version of your software. There are no other stuck/error states to consider.
I also implement reboot-with-no-screen-flicker, so that I can display a basic company logo so things look respectable while there is an outage and everything is bootlooping. Some brands of screen let you upload an image to display if there is no valid signal, which is an easy way to implement this.
Also worth doing a load test of your server to verify that it can withstand all your devices rebooting simultaneously and you won't suffer the thundering herd problem.
The goal is to prove forward progress and the best way to do that is to come as close to proving that your userplane SW is actually working and not dead, or worse, half dead.
Generally there’s a hardware watchdog implemented as a counter/timer in the processor. It can have a predefined or configurable period. It counts down, and if it times out then it initiates a hardware reset of the processor.
You can ensure your software/OS is always at least executing code by having a task (in-kernel on Linux, or an RTOS task, or just in your main event loop on baremetal) that resets that timer. Then, if your code stops resetting that timer, it expires and resets the processor.
Many x86 systems have a built in hardware watchdog.
Upon reading these I did realize on Linux it’s implemented as a kernel device, but it’s usually a userspace task that has to notify the kernel watchdog interface to actually kick the timer. This makes sense, since userspace being functional is probably what you really care about.
[1] https://www.kernel.org/doc/html/latest/watchdog/watchdog-api...
On most IPMI-capable BIOS/firmware there's now (been for 10 years but I'm old) an option to log 'system' events (ipmi failures like fan speeds if you've set threshold, but also reboot reasons). It's call the System Event Log. Very useful.
And on IPMI-plugged watchdogs, you can also see the state of the HW watchdog (is it running, how many seconds are left). Very useful too.
# /etc/systemd/system.conf.d/foobar.conf
[Manager]
RuntimeWatchdogSec=60
When used in this manner, if systemd fails to ping the watchdog for 60 seconds, the system resets.https://www.freedesktop.org/software/systemd/man/systemd-sys...
Somewhat related, nowadays by default systemd enables a 10-minute watchdog just before a regular reboot (i.e. after everything has been shut down) to ensure the reboot happens even if there is a hang for some kernel/HW reason.
[1] https://www.st.com/resource/en/reference_manual/rm0090-stm32...
It can be useful where you've got a mainloop doing some very predictably timed activities, and allows detection of faults which cause your watchdog servicing to occur too frequently.
I think it's quite common in DSP and things like motor control, where you often have hard realtime requirements and things happening too soon is just as bad as too late.
We use this sort of approach with diskless systems in particular. If there's a power cut, the first boot attempt after power restoration might not work (because the network isn't back up yet). So the diskless systems just sit there, continually attempting a network boot until successfu (at which point the software on the PC hits the 'restart timer' button periodicaly.
This is closely related to the concept of a "deadman's handle", for example train drivers who must keep a lever pressed down during operation - if it's released, the train stops automatically.
If the watchdog hasn't been fed for X milliseconds, then it resets the system.
https://shelly-api-docs.shelly.cloud/gen2/General/RPCProtoco...
https://www.amazon.com/stores/page/4E39C18F-DCA3-4726-A8A7-6...
Plus, they're also easy to reflash with other firmwares (e.g. Tasmota, ESPHome).
Edit: For small do-it-yourself computer clusters you probably want it to fail normally-open, this would be for non-redundant stuff
wmic RecoverOS set AutoReboot = True
Setting it to 60 seconds is not much use when it takes 90 seconds to capture the core file.
Use kernel parameter panic=60 in your kiosk
I was still curious about the content (which is how a title is supposed to work, right?) so I clicked it anyway, and learned about a kernel parameter I wasn't aware of.
It's a good article, thanks for writing it.
I wish there was support for this out of the box in linux distributions. Either by providing a (sub)set of software that doesn't need to write, and/or adding an overlay. Even for rasperries which are often used in a way where they would benefit it requires some work to set things up.
Having that kernel parameter set automatically when in that mode would be nice, too.
Or maybe I'm just not aware that distributions for this specific use-case exist.
https://blog.dustinkirkland.com/2012/08/introducing-overlayr...
Running as full XFCE Desktop with FF, occassional updates except kernel, all in RAM, with quick logout and telinit 1 as root from console, then back to telinit 5(with autologin back into X):
$ uptime 09:44:35 up 160 days, 46 min, 3 users, load average: 0.02, 0.08, 0.31
Most hassle-free Linux-distro I've ever had, while being compact, but fully usable for my needs. There may be others, but I'm tired of trying.
Edit: Obviously not running with this kernel-parameter or watchdog, because I don't see panics here. But could be easily integrated, because Debian. OFC you'd need to have a remastered and prepared(to your needs) boot-device available after a reboot, while running from RAM. Could even be an USB-keychain. Anyways, they have all the tools to do this within a few clicks from X and a few function key-presses and menu choices during first boot.
Works for me, YMMV.
I run into kernel panics a lot and usually power cycle the machine. In particular for this to reboot after the kernel dies it seems to need real hardware support. Ping this register every N seconds or I reboot style.
Is that what we're working with here, vs some other part of the kernel which might survive part of it dying?
> docs here suggest the system can get in a state which ignores the watchdog, https://mjmwired.net/kernel/Documentation/nmi_watchdog.txt
> thread here suggests a bug in the bios can stop a watchdog working https://community.intel.com/t5/Intel-NUCs/Issue-with-Linux-w...
Where ? The most I saw was when trying to boot something untypical.
> Is that what we're working with here, vs some other part of the kernel which might survive part of it dying?
I'd imagine it's same path for panic. There is probably still ways to crash (especially with hardware faults) to get around it so it isn't replacement for watchdog but it looks like an improvement
> docs here suggest the system can get in a state which ignores the watchdog, https://mjmwired.net/kernel/Documentation/nmi_watchdog.txt
I'd like to say you need "proper" external watchdog like IPMI based one on server but
* They still need to get thru the boot for kernel to activate them in the first place. I think I saw some servers that had some extra protection to crash on boot but definitely not anything common
* I've seen weird bugs around those too. Like older IBM servers could trigger watchdog when you synchronized IPMI time
You really kinda need 2 watchdog loops, long one (say 5-10 minutes) that runs from start catering to "something crashed during boot", and short one (few seconds to maybe a minute) that gets activated once kernel and OS is up and running
For example, do you have read-write filesystems or swap? If it gets corrupted, will a reboot cycle with fsck corrupt it further, preventing post-mortem forensics? This may have security implications too.
What about network access to a shared system? Will a power outage cause a stampede on your servers? Should you stagger that 60 value?
Just a few questions to ask.
The bigger thing is asking yourself these questions should reduce the pain.
When it doesn't, what are we doing?
vm.panic_on_oom = 2
In my experience, most systems get into a precarious state when not engineered to only use a percentage of available memory, factoring in the apps that spike memory using when forking child processes. This is most useful on NFS diskless machines, any machine that is managed as cattle vs kittens i.e. treating them like disposable VM's, dev/qa machines and especially machines used for application load testing.One could even use in production after a significant amount of testing but it requires a few more settings and having automation doing a little bit of math as OOM/vmscan often picks the wrong processes to terminate and leaves the machine in a broken and sticky state and increases downtime impact on SLA. The probability of an OOM panic can be reduced by using:
vm.overcommit_ratio = 0
The default of 50 or 50%, 1.5X ram+swap allows the kernel to permit an application to spawn even when there is no memory available. This is handy on development systems but can lead to problems when automation does not properly calculate how much memory is available and permits an application to request too much memory as the kernel does not protect it's own future memory requirements.One must then use a little bit of math based on the total memory installed in the system to calculate the values of vm.admin_reserve_kbytes, vm.user_reserve_kbytes, and most importantly vm.min_free_kbytes to prevent applications from grabbing too much memory. This also requires adjusting automation to deal with grabbing enough information when an application does not spawn correctly. This can be further improved by containerizing applications or using cgroup memory constraints when feasible. A percentage of memory must remain unallocated by applications for the kernel to use.
If using idrac/ilo/ilom cards, some of them have a feature to grab screenshots and will save the screenshots that led up to the panic, though will not always get the right moments.
[Edit] I should also add that the math calculations assume all deployment teams and all of their automation tools are aware of one another and math-all-the-things. e.g. production code, monitoring tools, analytic teams all add their numbers together plus overhead. If this is a Kiosk then maybe that is just one team.
Thank you.