Twenty Years of Valgrind (2022)
nnethercote.github.io
nnethercote.github.io
Also for leak detection, some IDEs have excellent realtime debugging tools (for instance there a very nice memory debugger in Visual Studio since VS2015, and XCode has Instruments).
Are there any use cases that people here have experienced where Valgrind is their first choice?
If I have anything that feels like UB, especially that I can trigger in a unit test, I run it through valgrind first. If the issue can only be triggered by running the whole system and valgrind slows it down too much, there's a good chance that sanitizers will also slow it down too much.
Among the sanitizers the valuable ones are the general undefined behavior ones. Valgrind won't tell you about your integer overflows and whatnot. Like that in b >> a, a turned out to be 32 or greater at run-time.
There are some situations where I find myself using Callgrind, in particular in situations where stack traces are hard to extract using a regular profiler. But overall, it's a tool that I find vastly overused.
In effectively all latency-sensitive contexts, sampling is worthless. 99.999999% of the time the program is waiting for IO, and then for a handful of microseconds there's a flurry of activity. That activity is the only part I care about and perf will effectively always miss it and never record it to completion.
I need to know the exact chain of events that leads to an object cache miss causing an allocation to occur, or exactly the conditions which led to a slow path branch, or which request handler is consistently forcing buffer resizes, etc.
I never need a profiler to tell me "memory allocation is slow" (which is what perf will give me). I know memory allocation is slow, I need to know why we're allocating memory.
Callgrind is a CPU simulator that can output a profile of that simulation. I guess it's semantics whether you want to call that a profiler or not, but my point is that you don't need a simulator+profiler combo when you can just use a profiler on its own.
(There are exceptions where the determinism of Callgrind can be useful, like if you're trying to benchmark a really tiny change and are fine with the bias from the simulation diverging from reality, or if you explicitly care about call count instead of time spent.)
In reality I really doubt they are.
Most Fortune 500 c suite are bean counters with abysmal engineering or product know how. They can’t see past the next quarter earnings report. I doubt long term contribution to meaningful open source is on their list.
these days, since I am doing bare metal embedded, I barely use valgrind; but it is game changing in the situations where it's useful.
(2022)
Some discussion then: https://news.ycombinator.com/item?id=32245136
"We implement our own memory allocator" is no excuse; the primitives you use should be hooked by valgrind so at most there should be false negatives due to your allocations being larger than the user-facing ones ...
There used to be one in LuaJIT because it had an optimized string comparison that compared outside of the allocation (which is allowed by the OS as long as you don't cross a page boundary, which LuaJIT's allocation algorithm made sure it never did)
The suppression was removed in https://github.com/LuaJIT/LuaJIT/commit/ff34b48ddd6f2b3bdd26... when the string hashing got a new implementation
But even for leaks, you can also have intentional leaks that valgrind will flag but that you can't really do anything about. One example is how using `putenv` can lead to you having to leak memory on purpose. There are many other cases.
For example the OCaml suppressions here are because OCaml (as expected) doesn't free static allocations at program exit. You can also see some real bugs we found:
https://gitlab.com/nbdkit/nbdkit/-/tree/master/valgrind?ref_...
I routinely run on 1TB memory 128 core racks, and I don't worry about free() much.
I'm not humblebragging, I actually think this is lazy (!) and I would benefit from more explicitly thinking about the memory consequences of what I do but there are some things which I used to freak out about growing to GB and now, I regard it as an investment on the future me, running the same thing: It's very likely I've got it in a hash structure of some kind already. I just add columns to the dict() elements.
Down the other end, I recall some friends getting code which I expected to have to run on a major rack host to build onto a small memory model rPi and they said rust did that: allowed them to get rid of the overhang of other languages expectations to runtime size and be explicit about use and free in the heap.
Lastly, the same fragmentation will cause advanced libraries like Eigen to use many small memory areas, and jumping between them kills locality, hence causing you performance losses.
I’m running tasks on (and administering) clusters with similar resources to yours, yet I always treat them like 486DXs with 4MB RAM, because I can’t restart them every week.
Now it could’ve been the case that some of those objects had a __del__ method that needed to be called or something. Absent that case, I’d have preferred the process just exited and been done with it.
If that were a program that ran synchronously from a shell script, the shutdown GC time would’ve been nearly as long as the data loading time. Maybe Python could benefit from a fast_shutdown_GC function that only calls free() if an object is something with a non-trivial delete method. Otherwise, skip it and let the OS do its voodoo it does so well.
I’m picking on Python here because that’s where I last saw this. The basic idea applies in lots of other cases though.
In fact I did this a couple of years back on a project I was working on maintaining in win32. It had a complicated memory management scheme for a GUI process that kept screwing up. Turned out it couldn't ever use more than 2MB of RAM so I just allocated 10MB up front at the start of the process and wrote a simple incrementing pointing counter in that allocated block and free'd it at the end like a short life pool. Solved all the issues. Machines it's running on all have at least 16GB of RAM so it's not exactly breaking the bank!
Even memory fragmentation is a problem for us. We once had a chat server that would crash every 3 months, which we tracked down to memory being fragmented by a less than optimal malloc implementation in glibc. (Since the glibc fix was in the "too hard" category, we reluctantly decided to schedule a restart every month.)
I moved to Java with the millennium, and missed the rise of Valgrind. I'm just happy that everyone has good tools now. Writing C in those days was a rough business.
Unable to find the problem, I ran the task on valgrind. Took almost a week, but it showed me the reason nice and clear.
Twenty years of Valgrind - https://news.ycombinator.com/item?id=32245136 - July 2022 (112 comments)
For verifying unsafe code, Rust has a MIRI interpreter that catches UB more precisely, e.g. it knows Rust's aliasing rules, precise object boundaries, and has access to lifetimes (they don't survive compilation).
Non-deliberate leaking of memory in Rust is not possible for the majority of Rust types. In safe Rust it requires a specific combination of a refcounted type that uses interior mutability which contains a type that makes the refcounted smart pointer recursive. Types that meet all three conditions at once are niche.
The only annoyance/incompatibility is that Valgrind complains that global variables are leaked. Rust does that intentionally, because static destructors have the same problem as SIOF[1] in reverse, plus tricky interactions with atexit mean there's no reliable way to destruct arbitrary globals.
The few times I used the cache simulator and compared it to (Linux) Perf I found reasonable results. Someone can recommend a better cache simulator?