My Favorite Debug Ever
ryan.hypnoticocelot.com
ryan.hypnoticocelot.com
As a second option I'm using valgrind which can find uninitialized memory (and doesn't require recompile of all libs as MemorySanitizer does).
If you're on Solaris or a Solaris-derived OS, you also have the libumem and watchmalloc libraries that can help you out. I have used libumem to great effect in the past.
It's been a while, but I anecdotally recall that Solaris is more stingy with its mallocs than Linux. I used to compile and run my C projects on Solaris as a first pass in my search for memory errors.
I guess it might be interesting for all the people who, like the author, are Java and PHP programmers and have no experience with C?
Maybe it isn't snark, but that reads like it, and it seems unconstructive. Why knock down someone who's learning by denigrating his background? What's the harm in someone sharing something interesting (to him) that he learned, whether or not you think he should have known it already? It's easy enough not to read it, especially on a random blog and even on HN, if you don't want to share his experience.
For example I am absolutely able to use gdb, but have never heard of Electric Fence. (But I had heard of AddressSanitizer mentioned above.)
It does have educational value: help to give more people pointers, short tutorial how to attack these issues. The more people can do this, the less bugs we will have around.
We need more people with these skills. It's all the time harder to hire these people nowadays.
Here's what I usually do in such cases:
Call stack is unreliable at that point. You'll see puzzling things like function calls with impossible parameters, etc. If things just don't make any sense, it's better to map all code paths that can lead to the crashing EIP/RIP (hopefully valid pointer to the instruction that caused the crash). Check EIP/RIP if it's in some rep movsd (= potentially inlined memcpy, check ECX (RCX) rep counter, EDI (RDI) rep pointer), or if the execution is in some runtime library code such as memset, memcpy, etc. similar. The next thing is to make sense of the call stack manually, if there are portions not overwritten, but what stack walk couldn't resolve. Of course it pays to take a look around stack otherwise as well, for signs of overwrite and contents of the overwrite.
It's also possible a pointer to stack object leaked at some point and the crash occurs at completely different part of code than where it actually segfaulted. Or some runtime structure was corrupted, like heap. Sometimes you can find those by just inspecting and guessing struct/object shape and values near stack pointer ESP (RSP).
If the bug can be reproduced, memory breakpoints, logging (especially if multithreaded, but watch out for blocking I/O from logging), tools for debugging memory corruption (valgrind, compiler paranoid mode, etc.) etc, even mapping some pages unreadable and unwritable. It can take a while to find the actual bug.
If reproduction is not possible, good luck. Better spend some quality time with memory hex view, disassembler, trying to locate registers and stack values that might contain pointers, etc. It might take a while to find the issue...
A next read in this vein might be Bug Hunter's Diary by Tobias Klein: https://www.nostarch.com/bughunter
The banking and healthcare industry already has this problem because there is a huge amount of decades old legacy stuff still in use and they have to pay through their noses for expertise.
Obviously, as you mention, the less there are people that are capable to do X, the more they cost. But it's not the same as "totally forgetting something".
Some low level people don't know how to do full-stack web development. Some full-stack web development people have no idea about memory management and stuff. etc...
You need all sorts of people to be productive.