The Azul Garbage Collector
infoq.com
infoq.com
* They are inserting special x86 instructions around object access JIT output instructions to "trap" uses of references to objects that had cleanup/relocation in progress. The trap works around the work in progress to "heal" the references in rare instances, usually allowing branch prediction on the x86 to simply fall through.
* It almost sounds like part of the relocation work by the GC used a mechanism not unlike "transactional memory" to either "commit" a block of moves of active objects, or roll them back in case of a conflict caused by the running application accessing/updating/creating something at an inopportune moment.
* One of the diagrams suggests that there are N GC threads corresponding to N application threads. If there is in fact a one-to-one correspondence, rather than just "there are many of both kinds of thread", I wonder if they have thread specific sub-heaps, and employ some kind of processor affinity binding together the application thread and its corresponding GC "shadow" on the same processor? Maybe that's automatic anyway, based on memory region in use? Anyway, localizing these tasks together might avoid processor cache misses. I may have read much more into one of the diagrams than was really meant, though. Even if they don't have thread specific heaps, I think I like the idea of having heaps tied to individual threads, only migrating objects/references to a global heap when they have in fact been shared between threads, or are anchored to some sort of static context.
Anybody care to provide an alternate interpretation of some of this?
^^^ Details how their read barrier works.
"you can't write memcached in Java" http://roboprogs.com/devel/2010.12.html
Or can you??? (efficiently, given the right JVM?)
At any rate, the defaults for Java are slower than those for Perl when doing many string operations on a single thread. Measurements: http://roboprogs.com/devel/2009.12.html. I have since rerun these tests on a 6 core AMD, with largely similar results. Of course, when doing threads or fork, the comparison breaks down as these constructs are implemented so differently between these languages, to say the least.
It's true that large heaps (multiple GBs) can cause long GC pause times (which is one of the problems Azul tried to solve). This can be mitigated by simply running more cache servers with smaller heaps.