Logging all C++ destructors, poor mans run-time tracing
raymii.org
raymii.org
Also valgrind, but I'm more familiar with the first two.
I think plain gdb would have been sufficient if it's exiting with a segfault or terminating...
You will find that in a typical C++ codebase, the destructors that do useful things (say, flushing useful buffers and closing files etc) are much fewer.
Then again, I guess it's optimal if you're writing fault-resistant software.
releasing resources (optimally) and resource exhaustion / ageing (crash recovery) are orthogonal.
> a destructor is the wrong place
I commented on GP's use of exit() not on finalizers / destructors.
> If an external service requires a client experiencing an error to manually release resources...
You perhaps missed what I wrote: It isn't optimal, if a failing client never signals release (think: distributed lock).
In concept. In the reality of implementation they are not.
> I commented on GP's use of exit() not on finalizers / destructors.
You said they were "suboptimal." Implying there is an available optimal solution. I'm challenging that exact notion.
> It isn't optimal,
You perhaps missed the point. It _can't_ possibly be optimal given the nature of the problem itself. These are the wrong terms to understand the problem in.
Recovery flows are very different in both concept and implementation.
> It can't possibly be optimal given the nature of the problem itself.
What's in the "nature" of this problem that one should never bother releasing resources once done? I mean, if any program creates an external resource (say, unix domain socket or shared mem), it isn't okay if it never releases it.
I still don't understand the hate for the C preprocessor. It enables doing this like this without any overhead. Set a flag and you get constructor/destructor logging and whatever else you want. Don't set it and you get the regular behavior. Zero overhead.
[You replace straight forwardly "mage" with "wizard" and oops, now your images are "iwizards" and your "magenta" is "wizardnta"]
There is no magic environment that can fix this for you. If you feel you've seen one, then you've focused on the parts of it that were important to you, while ignoring the parts of it everyone else actually needs.
Trace trace##__LINE__(__FUNC__);
The Trace instance would generate one log on construction and another on destruction. It also kept track of function call nesting (a counter) in a static member that would increment in the constructor and decrement in the destructor. It was inherently single-threaded, because I used a static member, but it could be adapted to multiple threads using thread local storage. I paired it with a LINE("Var x is " << x); macro for arbitrary ostreams-style logging. And building on that, EXPR(x) would do LINE(#x " = " << (x)). The output was along the lines of: ,- A::f()
| ,- A::g()
| | ,- B::B()
| | `- B::~B()
| | x = 12
| | About to do a thing...
| | ,- A::doAThing()
| | `- A::doAThing()
| `- A::g()
`- A::f()
The macros could be disabled (defined to do nothing) by a preprocessor symbol.How well does the visualizer handle multi-TB traces? Usually pretty uncommon, but a 10-100 GB is not that hard to produce when doing full tracing.
For the Bevy game engine, we automatically insert tracy spans for each ECS system. In practice, users can just compile with the tracy feature enabled, and get a rough but very usable overview of which part of their game is taking a long time on the CPU.
To be fair, you do still want some manual instrumentation to correlate higher level things, but full trace everywhere answers most questions. You also want to be able to manually suppress calls for small functions since that can be performance relevant or distorting, but the point is “default on, manual off” over “default off, manual on”.
The only way I can see doing this at compile time is with a compiler extension but then you are entirely locked in to 1 compiler.
Maybe if you compile with debug symbols but then well, you are shipping debug symbols...
If you want to really zoom, you need to get the hooks inlined and probably just written in straight assembly. You then need to optimize your binary format and recording system. You then need to start optimizing your memory bandwidth usage when that becomes the bottleneck. Your overhead in the end is basically limited by memory bandwidth; you can only shovel so many tens of GB/s of logging into memory. Note that persistence has likely been infeasible for the last 2 or so orders of magnitude; RAM is likely the only storage consistently fast enough for the data rates you want to generate when doing this.
We use sampling for the cases where this level of detail is needed as it has lower overhead.
What use case did you find this useful for?
You do need to allocate a ton of memory for the recording buffer to record sizable amounts of trace data. GB per core-second of trace or so (ring buffer so you get to see the last N seconds, not you need to run for less than N seconds) but that is fine during development on normal dev machines.
It is useful for everything. Why would you not want full traces for everything? It is amazing. We use it for everything internally where I work. Or rather, it is part of it. We actually prefer full time travel debugging during development and automated testing (again, overhead is low enough) but it is not available for everything. So sometimes we are stuck with just traces.
[2]: https://github.com/wolfpld/tracy/releases/latest/download/tr...
[3]: https://github.com/NixOS/nixpkgs/blob/nixos-24.05/pkgs/devel...