Debugging Firmware with GDB
interrupt.memfault.com
interrupt.memfault.com
Me: Customer is having their program X hit a SEGFAULT when they run it.
Manager: Did you transfer over a GDB version to backtrace the core dump?
Me: Yep.
Manager: What did it say?
Me: Well, when I run the same series of steps through the GDB version, it doesn't segfault, so I have nothing to backtrace.
Manager: Huh...
Me: What would you like me to do?
Manager: Move the production binary aside, drop in the GDB version. Leave it there.
The best way I've found to deal with this is:
1. Compile with -O2 (or for firmware, -Os) by default, and generate debug symbols in the elf file. Crank it up to -O3 only on a file-by-file basis with #pragma
2. Collect the process memory when an issue is hit
3. Debug it asynchronously using the debug symbols + the memory.
That way, the debugging & the running are decoupled and the timing is stable.
I'm in the camp that says all builds should at least be built with symbols, so you can still debug release builds. Maybe stripped out of the distributed version, keeping a private copy for coredumps from the field, if you're doing closed source dev.
Second, gdb can give you a stack trace with addresses if there are no symbols. But if you compiled with gdb, you can give pass the flags "-Wl,-M" to the link step, and it will spew a link map out stdout, which you can pipe somewhere. With that file, you can figure out what the addresses in the stack map correspond to.
If you really want to go hard core, once you figure out what function the problem is in, you can compile that file with "-Wa,-ahlms=<filename>.asm.out" to get an assembly listing for that file. From a bit of hex arithmetic, you can find the assembly instruction that corresponds to the crash. The hard part is correlating that back to the source code, since I don't know of a way in gdb to get that assembly output with the C/C++ source as comments.
objdump -S can do that.
If you run debug builds, the compiler fetches an uninitialized variable out of memory that the OS nicely zeroed before it gave it to the process. If you run even the simplest optimizer, it will say: "Well, obviously, no one cares what this variable starts with, because they didn't initialize it. So let's just use the garbage value left over in this free-at-the-moment register." Expunge your uninitialized variables, and then show me the disassembly code that the compiler got wrong. (Once upon a time I managed a piece of the validation for a C/C++/FORTRAN compiler suite. I have had this conversation more than once. :)
Exception: Real-time code with timing-specific hardware interactions. But that isn't an optimizer bug, that's a design issue.
The other 0.1% of failures kept my team busy enough. But I never allocated any time to your problem until you proved that you had no uninitialized variables.
OTOH, I'm wondering if GP is building with -Wall and has eliminated all warnings? At least that's what we're doing, and this yields good results.
I once had to add 7 NOP instructions at the beginning of the bootstrap/"BIOS" code I wrote for a Z80 clone. I couldn't understand why test programs seemed to crash[3] about a 15% of the time. Later, I discovered the same behavior in code that previously worked. I spent over two weeks trying to investigate, which only produced more confusion as the behavior would sometimes go away or get much worse randomly with each change I made.
I finally found the bug using a (hardware) logic analyzer to watch[4] what the CPU was doing on the memory buss, Something wasn't finished resetting inside the CPU. Any instructions that ran too early would trash the internal state of the CPU, causing later instructions to have problems like asserting multiple chip select pins. Multiple ROM/RAM chips would try to drive the buss, and everything dies. The instruction this happened on depended on which instructions were run while the CPU was still resetting. The NOPs simply delay startup to let the CPU stabilize.
[1] http://www.catb.org/~esr/jargon/html/H/heisenbug.html
[2] http://www.catb.org/~esr/jargon/html/M/mandelbug.html
[3] "crash" == CPU locked hard with no activity on the memory buss until /RESET was grounded by the watchdog timer (or me)
[4] with a 100-pin PQFP clip-on probe that wouldn't stay attached
That way the backtraces can be decoded. Although sometimes they are a bit butchered because it's a release build (we use the same config while developing and it's usable).
In this post, we’ll walk through setting up GDB for the following environment:
A Nordic nRF52840 development kit
Hoping this guide will be a useful reference for diving into the QMK[0] port to nRF52840 by Sekigon[1]. Scratch-built keyboards are my hobby and it would be nice to have something beefier than an atmega32u4 handling both the keymap/layout and bluetooth.I've written a number of these at my previous companies, which all get loaded by default when any developer was debugging (basically wrapping gdb in a 'make gdb' style call). That works for internal development at a company where everyone is using the same flow, but that's nearly impossible in the real world as every one has a different setup.
I'm assuming numerous companies have written a Linked List GDB Printer (such as https://github.com/chrisc11/debug-tips/blob/master/gdb/pytho...) or a script that prints useful information from the global variables particular to the RTOS running. These are all great, but they are a pain to install.
Is there really no better way to share these across the Internet other than "Copy / Paste this text into your .gdbinit"? I'm thinking it would be possible and relatively painless to share these scripts through PyPi, require them, and load them like a normal package, but I haven't seen this approach taken.
Being able to debug firmware is very important to me, up to the point that the fact a chip (or family) is supported by OpenOCD becomes a determining factor when choosing a component for a new project. Specially if the project is (or could be) open sourced.
I don't like proprietary debuggers. If a chip requires or forces me or my company to purchase Segger XYZ probe or a custom IDE, then it's not discarded right away, but back to the bottom of the pile.
When there is no alternative, I find myself writing OpenOCD flash drivers even before starting any FW development.
[1] https://segaretro.org/Sega_Mega_Drive/Palettes_and_CRAM#Back...
Pretty sure I optimized by putting a piece of tape at the exit line on the TV, then trying to make the color bar more narrow. :)