Everybody thinks about garbage collection the wrong way (2010)
blogs.msdn.com
blogs.msdn.com
The analog in a long-running program is an arena: it keeps allocating memory indefinitely, and then frees it all in one clump when the program is done with that phase of operation.
Multi-process languages like Erlang sometimes use this approach as a degenerate case of a stop & copy garbage collector. They start a process with a young generation heap sized large enough to contain all the memory the process will ever allocate (you can get reasonable estimates on this from previous runs), and so the GC never triggers and all the memory is freed in one block when the process terminates.
When I googled that, I found QT apparent has created some suppression files to help with using Valgrind. But I'm pretty sure that's recent and users struggled for a while with QT's approach.
http://developer.nokia.com/Community/Wiki/Using_valgrind_wit...
Which is a great way to kill that process, or arbitrary processes, when a previously well-behaved program is run on a system that's a bit strapped for RAM at the moment, or run on a somewhat larger-than-usual input, or both.
That, or bring the system to a standstill as the swapping subsystem flogs the disk nonstop.
Possibly first one, and then the other.
Reasonable resource consumption is an engineering decision based on context (target, expected problem size, etc).
Deliberate leaking is IMHO a totally viable engineering strategy. If anything, behaviour is actually more reliable and predictable than the alternatives apart from "upfront static allocation".
If you ran into problems and want the entire system to behave predictably, you could on Linux disable vm overcommit, set oom_adj (so leaker proc gets killed preferentially), set strict ulimits and ensure abort() on malloc failure. However if you have to go to these lengths, leaking might not be appropriate ;)
Think of compiled regular expressions for grep, or total counts for wc, or a rotating line buffer for tail, or the previous line for uniq.
It's quite possible to get pathological behavior even with GC or careful RAII, as anyone who's ever leaked memory in a Java program could tell you. Simulating a computer with infinite memory breaks down when the working set you're touching exceeds the physical RAM of the machine. Actually, leaking memory is the least of your concerns in that case, as leaked memory just gets paged out and doesn't cause any problems unless it exceeds the computer's address space.
If I am not mistaken, PHP scripts use precisely that.
And they can get away with it, because they are short-lived.
> but there are things you can do even if you have infinite memory that have visible effects on other programs
Moreover there are things you can do even without having these external resources that will mess up your program and that is just running the GC in a highly concurrent application. Except for systems with separate per/thread/actor heaps, when memory is allocated vigorously GC runs will start to affect your latency and parts of the program will get blocked and frozen.
As for releasing resources. In vital systems it is important to run watchdogs that release resources on behalf of a the main applications. Sometimes processes crash hard, segfault, and may not run their own finalizes. So these resource lists might have to be sent to a watchdog process whose only job is to monitor for the original program crashing and then releasing these resources.
I think he means "infinite amount of virtual memory". Virtual memory already tries to "simulate a computer with an infinite amount of (random access) memory". This makes me wonder if virtual memory can be used as a poor man's garbage collector?
Copying GC needs two same sized memory pools. One is used and one is empty. If you have same amount of virtual memory as ram, the empty unused one can reside in the virtual memory. When GC happens live data is copied (and compacted) from live pool to empty pool and the roles switched. With virtual memory, you can use as much ram as there is without affecting performance significantly.
If you do manual memory management, long running programs cause memory fragmentation. If allocated memory sizes are very variable in size, available memory can be half of the actual memory in the long run.
In theory, moving and compacting GC with virtual memory can almost double the available ram compared to manual memory management.
>The representation method outlined in section [http://mitpress.mit.edu/sicp/full-text/sicp/book/node118.htm...] solves the problem of implementing list structure, provided that we have an infinite amount of memory. With a real computer we will eventually run out of free space in which to construct new pairs.[http://mitpress.mit.edu/sicp/full-text/sicp/book/footnode.ht...] However, most of the pairs generated in a typical computation are used only to hold intermediate results. After these results are accessed, the pairs are no longer needed--they are garbage. For instance, the computation
(accumulate + 0 (filter odd? (enumerate-interval 0 n)))
constructs two lists: the enumeration and the result of filtering the enumeration. When the accumulation is complete, these lists are no longer needed, and the allocated memory can be reclaimed. If we can arrange to collect all the garbage periodically, and if this turns out to recycle memory at about the same rate at which we construct new pairs, we will have preserved the illusion that there is an infinite amount of memory.
A bit of a tangent, but I wonder GC is one reason Android devices moved to larger RAM sizes faster than iOS did. (iPhone 5 == 1GB, GS3/N4/One == 2GB.)
Under manual memory management or refcounting, running with RAM 90% full isn't slower than 50% full. (Your app code isn't slower, at least; who knows it if has effects via OS having less for cache or whatever.) But under GC, if 90% of RAM is full of live objects, you'll be forced to GC several times as often as you would if only 50% of RAM were full. So that 1GB->2GB bump might actually _more_ than halve how often you have to collect.
There could be other reasons for the diff--maybe part of it is Android's more permissive multitasking model, maybe non-GC-related memory-use differences between Android/Java and iOS/ObjC, maybe greater focus on power consumption from Apple or greater focus on specs from Android handset makers.
The only difficulty I can see is persuading the C++ constructor to run on an arbitrary address?
void* mem;
... snip ...
MyClass* thing = new (mem) MyClass;
And that doesn't allocate memory, but calls the constructor on the mem pointer. You are responsible for having allocated enough memory. You cannot call delete normally on this object though. You'll need to call the destructor manually as such: thing->~MyClass();
before the memory is deallocated. If the destructor doesn't actually do anything, you can safely skip this.I am blissfully unaware of almost all C++ specific syntax beyond basic class definitions, as I use it to add the odd syntactic construct to C (like operator overloads for + etc for my math 2- and 3-vectors) so I had no idea of this.
Also, the C++ FAQ http://www.parashift.com/c++-faq/ is useful for getting an overview on the corner cases of C++.
After which you should visit its inverse, the C++ FQA (frequently questioned answers, i.e. an anti-C++ answer to the FAQ) http://www.yosefk.com/c++fqa/ to learn why no-one in their right mind would ever use C++.
That won't teach you C++ per se, but would give you a useful look at the breadth of the language and how and why things happen the way they do.
vector<vector<T>> // wrong, uses a >> operator?!
vector<vector<T> > // compiles correctly.
In some ways I prefer what Objective C did, as a proper superset, and in others I'm drawn to projects like Cello[1]. I'm happy with using (and knowing only) a very small set of C++ features, especially where they're useful for defining simple algebras for objects! I'm sure I'll find another feature I like one day - this one (GGP) may be one of them.There is one OSS project I know that I never want to emulate: the code literally fails to compile via the python installer until you have run the makefile at least once, by which time it's finished generating all the header files. Whiskey Tango Foxtrot!
Although, if you read through it and subsequently decide that C++ is not the best language to use, I wouldn't call that the worst outcome....
It's not particularly effective a strategy for managing file handles unfortunately and if you forget to close your output streams, your program will crash with "too many open files".
I think the easiest way to think about it is, "finalizers don't prevent handle leaks, they sometimes clean up a handle leak that already happened".
I'm not sure if any debugger environments give a warning when a finalizer gets run due to GC but they ought to--it always indicates a programmer error.
myResource = foo();
try {
...
} finally {
myResource.release();
}
Some programmers (although I've never met one) try to avoid this boilerplate by writing a finalize() method on myResource that releases it. It's an anti-pattern. C# and Java have syntactic sugar now that makes the boilerplate shorter, which should help.