Memory Profiling: Introduction
easyperf.net
easyperf.net
Small question, what's the difference between (3) and (4)?
Instead I've added random sampling - e.g. if ptr % modulo > level - output it or not.
Another factor that slows down is doing a callstack capture. It's not for free at all, like on Windows it has to go through the exception handlers, etc. I think perfetto simply captures the whole stack (need to check again), and then offline decodes it - or something like this.
Windows ETW Tracing can also capture the stack, but I guess it'll incur also some penalty - it can't come for free.
I also wish there was some kind of standard binary format for emitting alloc/free sequences with callstack/etc.
A good alternative that also cuts down on the amount of data significantly is to strip away temporary allocations, say, only emit those allocations which were alive for at least X seconds.
Another thing: you can always just runtime attach and profile a partial time to reduce the amount of data recorded.
Finally: if your tools do so many allocations, maybe you should consider optimizing them...
Yes, in many languages you can combine a few things plus a core dump and figure leaks out too... but average people actually use MAT because it's largely transparent, and it can operate on running processes. Very few languages can compete with that in practice, much less with reasonable performance.
Shameless plug: not sure if I'd call it decent, but you might want to check my Bytehound for more in-depth analysis: