How to build highly-debuggable C++ binaries
dhashe.com
dhashe.com
if (breakpoint_1 && (x_id == 153827)) {
__asm("int $3");
}
No, don’t do it quite like that. Do: __asm(“int3\n\tnop”);
int3 is a “trap”, so gdb sees IP set to the instruction after int3: it’s literally the correct address plus one. gdb’s core architecture code is, to be polite, not very enlightened, and gdb does not understand that int3 works this way. So gdb may generate an incorrect backtrace, and I’ve even caught gdb completely failing to generate a trace at all in some cases. By adding the nop, IP + 1 is still inside the inline asm statement, which is definitely in the same basic block and even on the same line, and gdb is much happier.Add a breakpoint somewhere in the code, say added as breakpoint #2. Then;
condition 2 (x_id == 153827)
Or is there some other reason to not do this?Inserting an unconditional debug trap into a shipping, production executable, is a complete nonstarter. The program will receive a signal that is fatal, if unhandled.
Setting the build for an old x64 machine (https://dhashe.com/how-to-build-highly-debuggable-c-binaries...) for reversible / time travel debuggers seems unnecessarily restrictive to me. I'd expect a modern time travel debug tool (e.g. either rr or Undo - disclaimer, which I work on) to cope fine with most modern instructions (I believe GDB's built-in record / replay debugging tends to be further behind the curve on new CPU instructions - but if you're doing anything at scale it's not the right choice anyhow).
Regarding compilation (https://dhashe.com/how-to-build-highly-debuggable-c-binaries...) - we generally advise customers to use -Og rather than -O0. As the article states, this will still optimise out some code but should be a good trade-off without being too slow. (NB. last I checked, clang currently uses -Og as an alias for -O1, so it may behave less satisfactorily than under GCC).
It's also not said enough but: you don't need a special debug build to be able to debug. It's less user-friendly to debug a fully-optimised release build but it's totally possible. You just need to retain the DWARF debug info (instead of throwing it away). This is really important to know if you're debugging on a customer system or analysing a bug that's only in release builds.
Realistic scenario that gamedev uses: deoptimize translation units you are interested in finding or reproducing bugs.
Yep, looks like that’s this bullet point:
https://dhashe.com/category/blog.html#partition-your-tus-int...
This would be really unusual though right? For reference, in plain C code I see a 2x slowdown, in Zig a 4x slowdown (which I still need to investigate why exactly that's the case), and in C++ (even with heavy stdlib usage and on MSVC) at most 10x - which is the absolute worst case I've seen yet. My C++ info is a bit outdated though, have things gotten much worse in "modern" C++?
Or in other words: if you see a slowdown of 100x in debug mode, I would be really concerned about why the performance is so heavily dependent on the optimizer doing it's thing and would start investigating what's the reason for such a massive slowdown.
As far as I can tell, this is worse than for the average C++ codebase. I have some ideas for why the optimizer is able to achieve such large improvements with this particular code and some of that could surely be done with better source, but it's a huge code base and a refactoring on that scale just isn't going to happen.
Anyway the way I deal with it is to mark functions I want to debug with a macro that disables optimizations just for that function.
tracy_force_inline void swap( Vector& other )
{
uint8_t tmp[sizeof( Vector<T> )];
memcpy( (char*)tmp, &other, sizeof( Vector<T> ) );
memcpy( (char*)&other, this, sizeof( Vector<T> ) );
memcpy( (char*)this, tmp, sizeof( Vector<T> ) );
}
This is quite fast in debug.Will people who write C++ in 202x gamedev do it? Nah, I don't think so. Just stop even maintaining that pesky 'Debug' build and it will be fine. Build configurations grow in other direction. There are optimized configurations already that do not need local build step - LTO/PGO and friends. People just do not build that locally and real world performance is there.
YMMV of course.
If you've got a time traveling / reversible debugger than you can (sometimes) go back to a point where the value was being written / used, at which point it'll often reappear in scope and be accessible.
I believe DWARF's built-in virtual machine should be able to recompute missing values in many cases but I don't think compilers are great at putting the relevant info in, even where it should be possible to compute the right value fairly easily.
e.g. if the value you're interested in is being passed to / returned from a function then inspecting it around the call / return site should have the value available.
Pass various arguments to `backtrace` rather than just relying the default. Chances are it will have some non-optimized-out variables, which you can use to figure out what's going on.
Use `info registers` and see what looks like a pointer, then cast it to a type you suspect it is. Note that this can be done for any stack frame.
In that case, the code did become pretty hard to debug due to the extensive inlining and reordering it had allowed. Unfortunate because the only reason such functions exist is to make the structure of the code more apparent!
Maybe that's an exception (and / or maybe it's easier with C than for C++).
Maybe the less is that it's still always worth trying function call boundaries, in case the compiler has been conservative!
Note that this is highly dependent on choice of compiler. Clang is utter crap for debugging even at -O1, but I've encountered basically no trouble ever using GCC at -O2 (you do have to learn a little about how the binary changes but that's easy enough to pick up). I really would not recommend -O3; historically it introduced bugs and regardless it makes the build process much slower, and the performance gain is fairly negligible (I can't say how much it destroys debuggability due to lack of experience). I can't speak for MSVC personally but it's a bad sign that its culture strongly promotes separate debug builds.
That said, sanitizers are a place where a special debug build does help. Valgrind can do many of the things that sanitizers can but is around 10× slower which is a real pain if you can't isolate what you're targeting, so recompiling for sanitizers is a good idea.
(Other brief notes)
I have never actually encountered a case where the lack of frame pointers actually caused problems. As far as I'm concerned, any tool that breaks without them is a broken tool. (Theoretically they can speed up large traceback contexts if you're doing extensive profiling; good API design probably helps for the sanitizers case here)
Rather than assembly int3, Unix-portable `SIGTRAP` is very useful for breakpoints; debuggers handle it specially. You can ignore it for non-debugged runs but get breakpoints when you are debugging without changing the binary or options! Alternatively you could leave it unignored if you have tooling that dumps core or something nicely for you.
-O2: 10.22
-O3: 9.82
-O2 -march=native: 9.86
-O3 -march=native: 9.43
This is basically SHA256 over ~8GB of data, averaged over 5 runs. The numbers are rather crude here since I measured them just now, but I remember it was more significant when I first did it last month for https://news.ycombinator.com/item?id=40687942But - to anyone reading this later - please don’t do this blindly. You probably never want to distribute binaries with this flag set. It enables all the features available on the host CPU. So your build will change depending on the physical cpu you have installed. If you have a modern amd cpu, it may enable avx512 extensions and make your binary unusable on many Intel CPUs.
I've always felt that the debug code performance penalty was a good proxy for simulating what what users experience on machines that aren't god-level development machines. If it doesn't perform nicely on my machine with -O0, it's not likely to perform well on machines owned by mere mortals. And there's the extra lovely reward of being pleasantly surprised the snappy lively responsiveness of code that's compiled with -O3. (Optimizing actually-performance-critical code is of course, a separate kettle of fish).
This search is approximate both ways, but still finds a lot of examples of O3 being outright broken:
https://gcc.gnu.org/bugzilla/buglist.cgi?short_desc=O3&resol...
Of note, most of those (few) bugs tend to be related to specific language extensions.
> I believe GDB's built-in record / replay debugging tends to be further behind
Yep, I hit this issue on gdb’s builtin stuff. I added a footnote linking here and saying that this is probably unnecessary for rr and Undo. Thanks!
#define STACKTRACE GLOBAL_STACK_FILE[GLOBAL_STACK_IDX] = __FILE__; GLOBAL_STACK_LINE[GLOBAL_STACK_IDX++] = __LINE__; StackTraceCleaner stc;
StackTraceCleaner was a class that didn't do anything but execute GLOBAL_STACK_IDX-- in its destructor.So at any point in time I could inspect GLOBAL_STACK_FILE and GLOBAL_STACK_LINE and have a complete stack trace of the game.
Obviously this only worked because these games weren't performance-critical and because they were essentially single-threaded, but it did the job at the time. We're talking about a time when Visual Studio 6's support for templates was half-broken, and the STL wasn't exactly S, to the point that I had to roll out my own string, smart pointers, containers, etc -- made twice as hard because of the aforementioned broken template support in VS6 :(
I do miss these simpler, more innocent times, though.
STACKTRACE("e0957136-fed3-414d-80b9-8bbf84f3fa03");
With the GUID, I could see where functions moved as they were refactored.
I would write out the GUID to a thread-local file handle, along with a time-stamp, and an "enter" or "exit" when an RAII object left the stack.
Then I could retrospectively debug after running my program. I could see the callstack, and step through in time. I would walk my source and record GUID-to-filename/linenumber in a map. Then I could dump out a Visual Studio output that had the file name and line number, and execution time... allowing me to step forward and back through the execution.
Stone knives and bearskins.
I wrote a plugin for a past employer to visualize our internal product event hierarchy performance as if it were a normal C/C++ call stack, it was pretty cool. ETW and WPA are phenomenal tools. I miss them both dearly when on Linux.
Another techique is that all key allocations were handled based, so we could also easily dump what the whole process map was about.
During the translation phase, we inject a call between every 2 lines (of the original script language) similar in spirit to the macro above, that updates a global array with a counter. This global array is also built during the translation phase.
There's a separate thread which polls this array and sends updates to the IDE, so users can see in real-time what code was run and when, without needing to stop and debug the code or insert printing statement.
There's a somewhat complicated algorithm which during translation gives a unique ID to each callee in the call tree, so that thread sends to the IDE basically just a number with a counter. This doesn't deal with recursion BTW, I just limited reporting on recursion to a shallow depth. It still runs but just isn't reported.
It's not a complete replacement for a debugger (we also have a debugger) but it's good enough for most simple cases.
There's a similar hack for variables (including local variables).
Recently, some of my less technical users have been disappointed to discover this feature isn't present in other languages / IDEs...
You can also sort of do the same trick in any language, if compiled without optimizations the compiler will usually insert NOP assembly statements between each source line (disclaimer: not all compilers, not all languages, depends on a lot of factors).
This gives you both runtime visualization of running code and potentially time-travel debugging.
So one might be able to run a post-build patching phase and replace these NOPs with a call to a reporting function as above.
Out of curiosity I actually did this with C#, replacing NOPs in the MSIL level, since MSIL is easier to reason about than pure assembly, and it worked very nicely, I got a full "log" of program execution, including all lines executed, when, and all values of local and global variables.
I used the Cecil library to make it slightly more pleasant to read and manipulate IL.
Didn't go on with it other than writing a basic proof of concept.
There are of course tools which do the same for C++, like undo.io and RR, but I'm not aware of any tool doing this for C# / dotnet code. I'm not sure why since while not trivial, it isn't very difficult (and I'm not an expert by any means in this sort of thing).
Roslyn (C# compiler infrastructure) has code generators now, which is nicer than working with IL, but as far as I was able to see they don't support this scenario.
I'm really curious how this was presented to the user. A table with timestamp, filename, line number, and line contents? Or something more advanced?
Not a silver bullet but still, being able to collect and analyze user-mode crash dumps is sometimes the best way to investigate and fix bugs.
Thanks a lot!! I don't know how many times I've stepped into the C++ standard library and it gets really annoying..
At least on Windows you can setup Symbol Server + Source Indexing to achieve the same result.
Once upon a time I wrote a small tool that can embed full source code into PDBs. I doubt anyone has ever used it though. For proprietary software it's not uncommon to leak PDBs on accident at some point. It could be disastrous to also leak full source code!
https://www.forrestthewoods.com/blog/embedding-source-code-i...
It's relatively easy to add source indexing to PDBs. I've successfully done that for a non-standard Monorepo. Works great.
It builds a lot on quite a simple conceptual base, benefiting from native support in gcc / clang (for embedding unique build IDs) and in GDB (for contacting the server). It can serve up both source and symbol information.
I would like to see this adopted more - e.g. build infrastructure automatically populating a debuginfod server so debugging is seamless.
(There's no other out-of-the-box solution to this right? i.e. having symbol info live somewhere else other than the .so/exe, that can be loaded on demand when debugging? Like .pdbs basically.)
Some info here on how to configure GDB to use it: https://sourceware.org/gdb/current/onlinedocs/gdb.html/Separ...
The old-school way appears to be to extract the debug information from the binaries after compilation, then strip the binaries. As described here: https://stackoverflow.com/questions/866721/how-to-generate-g...
The new way is to use gcc's ability to generate split DWARF directly: https://interrupt.memfault.com/blog/dealing-with-large-symbo...
This will work with debuginfod but you don't have to have that running to use these - you can just supply the symbol directory when you want to debug.
Unlike the article we are using lldb rather than gdb ... and while I appreciate thats its possible _at all_ to script the debugger to do some pretty printing -- I found it quite a bit more frought to implement than initially expected ...
To take the Eigen example ... Eigen is a 'header only' library and offers Templated vector and matrix types. The types are template over (optionally), data type, number of rows, number of columns, and matrix row order (row major or column major). All that information is not actually even available at runtime -- just the type name (with the instantiated values for template arguments) ...
I ended up having to super hackily parse information out of the template type name in order to be able to pretty print the matrix appropriately in lldb ...
Problems of this nature abound when debugging c++ ... Very often with a header only library, there isn't even a symbol for methods you might want to call -- so you want to eg, call the size() method on some object within the debugger to see how big it is, you'll often be out of look due to an undefined symbol reference since the 0-overahead compilation models ensures the symbol doesn't even have to be created in the binary ...
Would be nice if there was some kind of way around that -- I guess I need to try the workaround mentioned in the article of explicitly instantiating template classes for common classes in 'debug' mode ... My fuzzy mental model derived from previous experience somehow doesn't think that will actually help the issue tho -- but I'd be happy to be wrong!
I’m using Windows and compile with Visual Studio. That debugger visualization file https://github.com/cdcseacave/Visual-Studio-Visualizers/blob... makes Eigen vectors and matrices show up nicely in debugger.
A nice trick with MSVC is you can turn off optimizations for TU or any block of code with:
#pragma optimize( "", off )
Leaps and bounds easier than hacking the build the system.I got pretty much the same on Linux when using CLion IDE from JetBrains.
#pragma GCC optimize ("O0")
(see also target, push_options, and pop_options)This is also available as a per-function attribute, using both gnu and standard syntaxes:
__attribute__((optimize("O0")))
[[gnu::optimize("O0")]]Well, except the Debug build in MSVC doesn't do half of the things from this list. Also, the list tells you how to use the compilation driver directly, so when comparing stuff, you would need to use the "cl.exe" compiler. For a "default" debugging experience it's enough to use CMake's Debug build type. It even has a built-in "release with debug info" build type.
Once upon a time it was widely said that ASAN should not be used for production code. The authors advocated against it and from a general-purpose security perspective it gives attackers a very large writable memory region at a fixed offset to play with. But over time I see more and more ASAN code in production on the theory that ASAN may make a system easier to exploit, but a memory corruption will make it easier to exploit. And so it's better to have knowledge of the issue.
Also, I've personally found the glibc malloc tunables very useful for debugging.
"just".
But generally, using ASAN in production is not what ASAN is for. If someone needs "memory safety" that ASAN provides, and doesn't care about slower runtime, then why did they use C++ in the first place? Just use Java. I understand this is not an option for old codebases though.
Also, using ASAN in production is like using a library for which the author states it's only meant for debugging and they doesn't really care about introducing any attack vectors in future versions. Even if it's not "exploitable" now, it might be in the future. Why would anyone want to use such library and take the responsibility that nothing bad will happen on the customer machine?
If your call graph has more roots than your neighbors garden and the whole thing is a forest and not a tree you will have a hard time understanding, analyzing and ultimately debugging.
I guess I'm confused by the mention of call graph roots; in my mind those are just ... entry points. Edges in the call graph might be a PITA to follow in a debugger because of indirections like vtables/optimizer inlining/etc., but isn't that separate? Or is my terminology model wrong?
But specifically for runtime debugging, you can get the actual stack trace, so runtime polymorphism is not much an issue. In fact a debugger can be a convenient way to understand how complex class hierarchies end up interacting.
I think there's a huge amount of complexity both inherent to the problem and caused by fifty years of accumulated bad habits, which is indicated by the thousands of lines of code in compiler-rt dedicated to handling this issue. I'd like to call their library functions but they are all in private-looking namespaces. I also tried to use the Abseil failure signal handler but it often fails to unwind and even when it does unwind has a habit of just printing unknown for the symbol name or file, and never prints the DSO base addresses.
Or, indeed, getting a core dump and applying GDB to it. GDB seems generally pretty good at reconstructing stacks at arbitrary points in application runtime.
We've also used a combination of libunwind and https://linux.die.net/man/1/addr2line to produce good crash dumps when GDB is not necessarily available.
ETA: Thinking about it, I'm not really sure what it'd do for C++ - I guess you'd end up with mangled names, so if you want sensible names you might need to demangle (either as a post-processing step or within the dumper) too.
I don't think you'll get any decoded argument values out of it either, so I guess it depends what backtrace info is needed.
You can use `dl_iterate_phdr` at startup if you need DSO info?
As for whether or not you can use this in a signal handler... well, I hate reading the POSIX standard with regard to signal safety because it's just not well-written, but as far as I can tell, a non-async-signal-safe function can be safely called from a signal handler for a synchronous signal (which most of the interesting signals for dumping stack traces are--it's only something like dump-stack-trace-on-SIGUSR1 that's actually going to be an asynchronous signal), so long as it is not interrupting a non-async-signal-safe function. So as long as you're not crashing in libc, it should be kosher.
If you have timestamped module load/unload info with base address + range, plus context switch times that allow you to figure out which specific thread & address space was running at any given CPU node ID + point in time, you can always answer that question. (Assuming the debug infrastructure is robust enough to map any given IP to one specific function, which it should be able to do, even if the optimizer has hoisted out cold paths into separate, non-contiguous areas.)
I realize this isn't very helpful to you on Linux (if it's any consolation I'm on Linux these days too), but, sometimes it's interesting to know how other platforms handle it.
>Compile with frame-pointers.
It's good to see that enabling frame pointers are included in the recommendations for debugging purposes.
The discussions on the relevance and the usefulness of frame pointers earlier this year on HN [1]:
[1] The return of the frame pointers:
> If using libc++:
> Add this define to your CXXFLAGS: -D_LIBCPP_HARDENING_MODE=_LIBCPP_HARDENING_MODE_DEBUG
... until the bikeshedding bastards change the define yet again in the next release.