Memory leak proof every C program
flak.tedunangst.com
flak.tedunangst.com
> It is [...] entirely optional to call free. If you don’t call free, memory usage will increase over time, but technically, it’s not a leak. As an optimization, you may choose to call free to reduce memory, but again, strictly optional.
This is beautiful! Unless your program is long-running, there's no point of ever calling free in your C programs. The system will free the memory for you when the program ends. Free-less C programming is an exhilarating experience that I recommend to anybody.
In the rare cases where your program needs to be long-running, leaks may become a real problem. In that case, it's better to write a long-running shell script that calls a pipeline of elegant free-less C programs.
If you write a C library, it is a good practice to leave the allocations to the library user, or at least provide a way to override the library's allocator. Allowing your user to write a free-less *program*.
Which means that if you're writing a program, you probably are also writing one or more libraries.
I also know someone who maintains some Clinton era encryption. that code is controlled who can know about it as obsecurity was all you were allowed. There are other pathalogical cases where you can't know everything about your program-
Extending something you did not originally write may be quite different, of course.
2nd approach for when you ignore the first) you use a type that doesn't "own" the pointer, and have the FFI side that allocated it free it
3rd approach for when you find out you can't do the 2nd) hope that the "owned" type is actually generic over the allocator used, and make an allocator that doesn't free (or even better, calls the correct free over the FFI boundary). `allocator-api` is the effort in this direction.
This is true irrespective of language.
#define free(_) /* no-op */That does it. I am not going to use it.
#define free(_) ((void)0) #define free(e) ((void)(e))But I remember the first time I saw such a program which never freed anything: jitterbug, the simple bug tracker which ran as a CGI script.
It indeed allows a very simple style!
Meanwhile, use ccan/tal (https://github.com/rustyrussell/ccan/blob/master/ccan/tal/_i...) and be happy :)
https://devblogs.microsoft.com/oldnewthing/20180228-00/?p=98...
What happens when it runs out of stack? It just resets the stack pointer back to the start.
1. The module that failed actually was no longer in use for that part of the flight.
2. The acceleration for the part of the flight where it was in use was within the parameters.
3. Instead of just ignoring / clamping the out-of-bounds values, the no-longer-needed module sent a big error dump.
4. That error dump was sent to the next module in-line, which was expecting...I think numbers, but certainly not big textual error dumps. And so that failed as well.
5. I don't know if that second module was needed.
So had the module just (a) processed the incoming data normally or (b) clamped the values silently or (c) dropped the out-of-bounds values silently, the Ariane 5 would not have exploded.
Instead it did the "make it impossible to represent invalid states"-thing and exploded.
While I understand the appeal of that idea, I think it is overrated with any software that has to interact in some way shape or form with the real world, however indirectly. Because programmers tend to have a pretty limited understanding of what states are valid or invalid in the real world.
(See also: the "falsehoods programmers believe about XXX" series)
Then you just say terrorists were hiding inside…
If it's a fixed function synth, sometimes just allocating everything up front makes more sense.
It was a deliberate design choice that was surprisingly close to what the original article describes.
We found the underlying issue that caused memory pressure in the Gen2 region but the fix was to change some very fundamental aspects of the service and would need to have some significant refactoring. Since this was a legacy service (.net framework) that we were refactoring anyway to run in new .NET (5+), we decided to ignore the issue.
Instead we adjusted the GC to just never do the expensive Gen2 collections (GCLatencyMode) and moved the service to run on higher memory VMs. It would hit OOM every 3 days or so, so we just set instances to auto-restart once a day.
Then 1 year later we deployed the replacement for the legacy service and the problem was solved.
As an example, constant data that is allocated by GNU Nano is never freed. AFAIK, the same happens when you use GTK or QT; there were even tips on how to suppress valgrind warnings when using such libs.
[I love golang, and I think it's one of the best languages around. If only it had a truly optional garbage collector. But then again, it wouldn't be go I guess...]
In short: If it works until it crashes, it doesn't work.
Yet, the idea of not freeing some memory in your program is not entirely stupid. Unless your memory is allocated inside a loop of unpredictable length, it's not really necessary to ever free it. Worse: the call to "free" may even fall after your program has run successfully. Thus, avoiding the useless (but typical) freeing spree at the end of your program may make it more robust!
Where did this idea come from? I have seen leaks where a program can consume all available memory in just a few seconds because the programmer (definitely not me...) forgot to free something in a function called millions of times.
Counterpoint is that debugging leaks is ~hopeless unless you have the ability to prune “intentional leaks” at exit
Not in general. It depends on your debugger. For example, valgrind distinguishes between harmless "visible leaks", memory blocks allocated from main or on global variables, and "true leaks" that you cannot free anymore. The first ones are given a simple warning by the leak detector, while the true leaks are actual errors.
... or unless someone decides to convert it into a daemon "because of that ticket" and then QA goes all "oh, ah, the routing is dead, the sshd is dead and the whole box is all but bricked, what could've possibly caused that".
The function just indirects malloc with a wrapper so that all of the memory is traversable by the "bigbucket" structure. Memory is still leaked in the sense that any unfreed data will continue to consume heap memory and will still be inaccessible to the code unless the application does something with the "bigbucket" structure--which it can't safely do (see below).
There is no corresponding free() call, so data put into the "bigbucket" structure is never removed, even when the memory allocation is freed by the application. This, by definition, is a leak, which is ironic.
In an application that does a lot of allocations, the "bigbucket" structure could exhaust the heap even though there are zero memory leaks in the code. Consider the program:
int main(int argc, char** argv) {
for (long i = 0; i < 1000000; i++) {
void *foo = malloc(sizeof(char) * 1024);
free(foo);
}
return 0;
}
At the end of the million iterations, there will be zero allocated memory, but the "bigbucket" structure will have a million entries (8MB of wasted heap space on a 64-bit computer). And every pointer to allocated memory in the "bigbucket" structure is pointing to a memory address previously freed so now points to a completely undefined location--possibly in the middle of some memory block allocated later.There are already tools to identify memory leaks, such as LeakSanitiser https://clang.llvm.org/docs/LeakSanitizer.html. Use those instead.
Clearly the author of TFA is aware of such tools, since the idea is to trick them.
You know this isn't serious right?
While the provided code may seem like an interesting approach, it's important to note that it introduces a number of issues and potential pitfalls. This code is an attempt to intercept the malloc function using the dlsym function from the dlfcn.h library and store every allocated pointer in a linked list called bigbucket. However, there are several problems with this solution:
Portability: This code relies on the dynamic linking functionality provided by the operating system. It may not work on all systems or with all compilers.
Concurrency Issues: This solution is not thread-safe. If the program uses multiple threads, concurrent calls to malloc may result in race conditions and data corruption in the bigbucket linked list.
Incomplete Solution: This code only intercepts calls to malloc. If the program uses other memory allocation functions like calloc, realloc, or custom memory allocators, memory leaks may still occur.
Performance Overhead: The code introduces additional overhead for every memory allocation, potentially affecting the program's performance.
Undefined Behavior: Overriding standard library functions like malloc can lead to undefined behavior. The behavior of the program is no longer guaranteed to be consistent across different platforms or even different runs.
Limited Practicality: While this approach technically prevents memory leaks by keeping track of all allocated pointers, it does not address the root cause of memory leaks, which is the failure to deallocate memory when it is no longer needed. Encouraging developers not to free memory is not a good practice and can lead to inefficient memory usage.
A better approach to avoiding memory leaks is to adopt good programming practices, such as carefully managing memory allocation and deallocation, using automated tools like static analyzers and memory debuggers, and, when applicable, leveraging programming languages with automatic memory management (e.g., garbage collection in languages like Java or Python).Can you elaborate on that?
If it requires you to free correctly, then you can just use plain C and the smart pointer isn't accomplishing anything. If it doesn't require that, how does it work?
Traditionally memory is considered "leaked" if it is still allocated but nothing point to it; i.e. there's no way to navigate to the allocation anymore.
He has made a joke "solution" by simply permanently storing a second pointer to all allocations so that by this definition they never technically leak. You can still always navigate to ever allocation so no allocation has leaked.
Of course it's not a real solution because it doesn't actually change the memory characteristics of a leaky program; it just hides the leak. In other words the technical description of the leak above isn't really the thing we care about.
Seems like almost nobody here got that.
"If you don’t call free, memory usage will increase over time, but technically, it’s not a leak."
Define "it" here.
Because "just don't free" is pretty different from what's in the post!
The joke and the sometimes valid solution are different strategies.
The joke only reminds people of the valid solution. It feels like multiple people are giving the article the briefest skim and assuming the two are the same. Perhaps focusing on the words "optional to call free" and extrapolating off that without looking at the code or properly reading the rest of the text.
Unless I'm badly missing something?
I agree, and would expand the idea to all kinds of resources (files for example). Sadly not many languages have "automatic resource management". For example Go has automatic memory management, but if you read an HTTP request's body you have to remember to call req.Body.Close(). If you open a file you have to call file.Close(). If you launch a goroutine, you have to think about when it's going to end.
I'd like to know if some languages manage to automatically managed resources, and how they do it.
VB Classic has reference counting and the terminate event is fired when the count goes to zero. So as long as you don't store the reference in a global the terminate event is guaranteed to run when it goes out of scope (so long as you don't have circular references of course).
C++ has Resource acquisition is initialization (RAII)
We have a counter that goes up by 1 every time you call malloc.
And down by one every time you call free.
And when the program quits, if the counter isn't zero, an email is fired off and a dollar gets sent from the developers bank account to the users bank account...
And it is not even always appropriate. It is common to allocate some memory for the entire lifetime of the process. For example, if your app is GUI-based and has a main window, there is no need to free the resources tied to the main window, because closing it means quitting the app which will cause all memory to be reclaimed by the OS. You can properly free your memory but it will only make quitting slower. Usually programmers only do that to satisfy leak detection tools, and if the overhead is significant, it may only be done in debug mode.
Isn’t that close to how ARC (automatic reference counting) works?
static void *(*nextmalloc)(size_t) = NULL;
if (!nextmalloc)
nextmalloc = dlsym(RTLD_NEXT, "malloc");
}
Somehow the fact that the optimization is incorrectly missed here feels appropriate ;-)the other bug of course is that it's not thread safe
#ifdef HAVE_VALGRIND_VALGRIND_H
if (RUNNING_ON_VALGRIND)
#endif
free_all()
free is way too slow if not needed, so detect valgrind via its API. Just on valgrind do the unnecessary free dance. ASAN's memleak detector is disabled via its env.Perl5 does its final destruction similarly, only when it has important destructors (like IO, DB handles and such) to call.
A better way is to have an always existing inline running_on_valgrind() function, and use the #ifdef only for that function definition, either within it or around it (having it around the function also allows it to be inline only for the trivial not-defined case). Examples of this "inline function" style are found all over the Linux kernel (which has lots of conditionally-compiled code).
In terms of memory management, PHP has both reference counting and garbage collection. Internally, an arena per request is used[0], so leaking memory over long periods of time is fairly rare and usually limited to native extensions.
[0]: https://www.phpinternalsbook.com/php7/memory_management/zend...
So for some definition of “process” it holds regardless - it’s just not always a real OS process.
They added GC at some point to allow scripts to run longer (eg for websocket servers) but even then last I checked you had to manually enable it.
It used to be a quite common approach in UNIX server programming.
OS handles and such are still an issue, but you can wrap them too. Thank goodness for RAII, a fantastic reason to use C++ even if you're otherwise writing C-like code.
Memory being "reachable" is a property used by garbage collectors to determine what can be safely freed, but reachable memory can still be a memory leak. (Which is why languages with GC can still suffer from memory leaks...)
Our Salesforce implementation consultants had put 500 lines of 'x = 1' into a piece of code to force it to deploy -- and these were people at a top consulting company with a very lucrative hourly rate.
No idea if SFDC still works this way or if this would fly today, this was back in the days when you had to use Flex to integrate anything with the UI.
In any other scenario, if some execution thread reaches the point where it returns from calling lock() on the mutex guarding malloc(), it must eventually reach the call to unlock().
In any other scenario, if some execution thread reaches the point where it returns from calling lock() on the mutex guarding malloc(), it must eventually reach the call to unlock().
Over the years I had been debugging multiple memory problems and only once I had found a true memory leak. It wasn't the bug I was looking for, accounting for only 8 bytes per hour and it was the easiest of all: directly discovered by valgrind.
All others were memory bloat cases: memory reachable but not used. Like forgetting some small per-request descriptor in some auxiliary hashmap in one of the code pathways. Reproducing and debugging was pain because it required running a loaded backend for days under a profiler, before the bloated part becomes visible.
Funny thing is that while C++ is notorious for memory leaks this particular kind of problem is possible in any language that allows a global mutable state (so all the practical ones). From experience it seems that true leaks are unlikely in a reasonably-good C++ code. Just never mix business logic with memory management and you're good to go.
Memory hoarding. I hate it but I love it.
Technically all pointers are accessible, but the issue remains that we have not logically accounted for unused resources and are wasting capacity. In this sense, memory leaks are possible in all languages. We could store every object into a global structure, and it will never be freed in any language. Thus, my Rust program balloons forever.
I get the feeling that this might have been response to recent "memory leak" thread(s): https://news.ycombinator.com/item?id=39041520
After running your test suite, it kicks into action and deletes all the lines of your code that were never executed.
[1] https://fgiesen.wordpress.com/2012/04/08/metaprogramming-for...
I have written medium to large projects using this approach and a leak-free program is not outside the realm of possible
Hopefully the invariants are documented, but relying on documentation to help programmers avoid memory leaks is far from foolproof.
Now if we could just fix off-by-one errors.