A Practical Guide to Watchdogs for Embedded Systems
interrupt.memfault.com
interrupt.memfault.com
Doing some basic groundwork and implementing the timer/software watchdog, writing out to logs when a watchdog is about to trigger, and not disabling the watchdog during development are all basic things teams can do to more easily catch these issues before shipping firmware.
JSYK: I do believe this was mentioned in the post as well, under the section 'Adding a “Software” Watchdog', but when a little bit further and implemented a per-task watchdog.
I implemented something similar on a particularly annoying 8/16-bit controller just a few weeks ago. Extra fun since it had no instruction to read the program counter (and no general purpose register wide enough to hold it).
I hope ARM adds that (or IC vendors like ST if possible) to the core itself, just mirror the PC to a register on reset. Should be trivial in hardware.
1: https://mcuoneclipse.com/2015/04/04/poor-mans-trace-free-of-...
Fun story: I've only had this fail on me once. There was a hardware bug in the interrupt controller which caused the wrong vector to be invoked for NMIs.
Anyway turns out I really needed a full PC hardware watchdog. I ended up buying some $8 anonymous piece of Chinese hardware that's USB powered. It hits the motherboard reset switch if the motherboard hard drive activity light hasn't flashed in awhile. Dumb thing, but it seems to work.
The names are a bit weird to my taste. FreeBSD ones sound a little bit better (but maybe I am just more used to them): ichwd (ICH WD) and amdsbwd (AMD SouthBridge WD).
An advanced topic is enabling support for the watchdog in your bootloader and having a defined recovery path when the system fails to load or, worse, the application falls into a boot loop.
If you have the space, you can fall back to a recovery image or duplicate of the application. If you don’t have the space, falling into a DFU mode is a good plan.
Once you get the hang of it, it becomes pretty straightforward. And it's great training on how the HW looks like and functions!
Granted, on more lower powered systems, with fewer, more primitive debugging options (looking at you, PIC) things miiight look a bit more painful. Thankfully, such systems are (much) smaller.
IMO, the _most_ painful issues to debug are memory corruption issues happening on large systems without MPU enabled.
Collecting enough state to fix them from a customer report is tricky, I would say.
In smaller systems for memory checks I always have a complete self-check routine at boot. Although I cannot even remember when I had the last error for internal memory. The quality seems to be quite high, even after many read/write cycles. Never implemented that for a system with virtual memory.
What I mostly miss in embedded systems isn't debugging, it is having real unit tests without having the target board and all the periphery installed in a laboratory, which is the only location I can feasibly run any tests. And those are mostly restricted to integration tests.
What's less fun is if there is too little protection against electrostatic fields/EMI on the JTAG clock pin. On the small cortex m-class devices we work with, some of them can't shut off the JTAG part of the chip, meaning that when operating, if there are enough (I think 8) logic flips on the TCK pin in _any_ amount of time, the JTAG part wakes up, sets the HALT ON BOOT flag. Next time the device reboots (due to firmware update, or watchdog, ...), it will stop and stay in JTAG debug mode. Not nice. You need to manually power cycle the thing.
We detect this by periodically checking the JTAG power domain, and if it is on, tell the server this so that we avoid rebooting it (eg automatically after firmware update). This way we've found poor hw implementations and tough EMI environments by proxy of JTAG power domain :D.
TCK has usually a pull-up/termination (but I've also seen pull-downs with/without caps). You see this issue even with the pull-up/down?
We now see this primarily on 1st/early iteration hw and on prototypes, but checking this proactively have saved us from a lot of headaches of the type "why doesn't it come back online - hw, or something in the update, or what...".
Some mcu pins also go into an unknown state (neither guaranteed high nor low), so resetting a cpu can have bad consequences if it's driving big machinery, if not designed correctly.
One project I had a pc sending software watchdog pings to several independent devices and each of those had an actual hardware watchdog (as opposed to the cpu resetting one in the article). I used the watchdog to physically control the power to contactors: no watchdog = no power = nothing activates.
The system controlled firing of gas burners and fans etc, but the design was very safe, heaps of redundancy and was guaranteed to fail into a safe mode at any instant.
Not in America but terms I've heard include feeding the watchdog, keeping it alive, petting, greeting, barking at it, calling it and, my favourite, shushing it.