New release of Memory Pool System Garbage Collector
mailman.ravenbrook.com
mailman.ravenbrook.com
> The MPS has been in development since 1994 and deployed in successful commercial products since 1997. Bugs are almost unknown.
Then half way down the page is a list of known bugs.
MPS appears to have been designed in a single core era, and it would be great if you would address any architectural changes that were made to address the current prevailing multi-core platforms.
/tia
[edit: removed ref to the azul q]
Most of the problems are around the bottleneck of stopping threads on other cores in order to get privileged access to memory while preserving their consistent view of the heap. What the MPS needs is its own protection map. Most OSs don't help much with that, although Mac OS X, for a while, allowed you to map the same physical RAM at two addresses with different protections. Sadly, no longer.
The restriction could be removed if either:
* the MPS could use a different set of protections to the mutator program
* the mutator program uses a software barrier
The first one is tricky, and the second one just hasn't come up in any implementation we've been asked to make yet. Given a VM, it could happen, and the MPS would be concurrent.So, I believe there's nothing fundamentally non-concurrent about the MPS design. It's kind of waiting to happen.
[I've put this in a comment in the code. Thanks for making me write it down :)]
e.g. how does it compare with the Boehm collector?
And hopefully this will change forms in the coming months to something more modern / up to date.
There's also some stuff here: http://www.ravenbrook.com/project/mps/master/manual/wiki/
MPS allows for precise rather than conservative GC. It supports doing copying collectors (see poolamc in the design docs for example).
It has some pretty nice telemetry and event logging as well.
Overall, I really like MPS and it is pretty powerful. It is a nicely structured library. It is more work to use than Boehm, but it does more than Boehm.
If the licensing terms aren't suitable for you (for
example, you're developing a closed-source commercial
product or a compiler run-time system) you can easily
license the MPS under different terms from Ravenbrook.
Please write to us <mps-questions@ravenbrook.com> for
more information.
(And you can see from the exception that we got for OpenDylan that such things are possible.)I was excited to use MPS in the runtime of an open-source compiler, but if users of the compiler had to pay in order to use it for closed-source apps then that would be a big turn-off.
Additionally with CMS, there is global heap fragmentation and recompaction which can take literally minutes on a large (8gb+) heap or slow machine.
Setting aside exact figures, can you comment on the above challenges and issues?
That said, we have an abstract framework that could be made to work in other ways, depending on requirements. It wouldn't be much work to hook in a write-barrier-only pool class, etc. etc. and have it co-operate. However, most of the development effort so far has gone on the read barrier approach.
As to how this effects overall run-time, well, that's where we'd have to arrange a side-by-side comparison, and make sure it included one of those compactions :P
- multicore/multithreaded - heap sizes of 20-200GB and larger ideally
Think about it this way... We have the same goal, of diminishing the use of manually allocated ram to the smallest possible place. Think of the GC vs malloc as compilers vs assembly language. Assembly has it's place, but it's no longer because compilers cannot generate efficient object code in a vast number of circumstances. Lets do the same for GC!
And we're definitely not proselytizing garbage collection here. The MPS is a framework for both manual and automatic memory management (and co-operation between the two). One of our main high performance commercial applications is all about the manual management, and for very good reasons.
But nobody should be rejecting GC out of hand, that's for sure.
For those of us too lazy to dive into the code, can you say what this read barrier is?
Since we're just moving our commercial clients onto 64-bit, we don't have experience except up to the 4GB limit (or 3GB in old Windows) which they approach regularly. We're still waiting for it to become a problem. I may have a different opinion in a year.
One of my short term goals is just to get some measurements together to publish. Watch this space, and sorry that they aren't here yet.