The Managed Runtime Initiative
lwn.net
lwn.net
Where a reply to this posting is made. Further comments following are generally quite interesting, e.g.:
"In general, current virtual memory implementations in almost all OSs assume that virtual memory manipulation is a relatively rare event, and we have put forward an algorithm and an application for rapidly-changing mappings that makes a real difference to a vast array of applications, but it can't do so within the limitations of current virtual memory APIs and manipulation speeds."
"[...] Our GC code is all in user space (part of the OpenJDK based code we put up along with the kernel mods) - it just needs some very scalable and somewhat different virtual and physical memory manipulation semantics form the kernel." (http://lwn.net/Articles/392745/)
And:
"The duration of the stop-the-world pause in all current JVM GC's is generally linear to the amount of live data the heap contains (you have to scan all that stuff and fix all the pointers to the relocated objects). This means that the larger the heap - the larger the pause. Sun's CMS (the Mostly Concurrent Mark Sweep -XX:+UseConcMarkSweepGC mentioned above) will delay the compaction as long as it can and track empty spaces in free lists, but it will eventually fall back on it's compaction code and pause for about 2-4 seconds per live gigabyte on a modern x86-64 machine. This is why JVMs are generally not used with more than a few GB of data, except for batch apps (ones that can accept a 10s of seconds of complete pause). Since a 256GB server now costs less than $18K, there is a ~100x and growing gap between commodity server capacity and the ability for individual runtime to scale with acceptable response times." (http://lwn.net/Articles/392797/)
Azul has claimed they've fixed this on their own custom hardware (a generic 64 bit RISC with some extra instructions including at least one to implement a fine grained memory barrier) and this is a version of their software based on that.
It will be very interesting to see if it's truly practical to put all your GC code in userspace; they made a decision to put almost all of their secret sauces in userspace and say that's saved them many times.
I write code in Clojure, deploy on the JVM using Amazon's Linux servers and GC pauses are a very real issue for us. Azul has been working on this problem for years and I'm quite happy that they decided to open-source a large part of their work.
Now, before you jump on me with the obvious — I do not advocate pushing crappy or unclear code into the kernel. But I'd much rather see a healthy discussion than this childish "take your toys and go away, we don't like you and you're not welcome here" attitude.
Given that you run clojure, have you considered trying to use a concurrent garbage collector?
It's a hard problem and especially so on stock hardware where doing the sorts of things they do is necessary, e.g. read barriers implemented with virtual memory hardware and a fast interface to that.
I've considered playing with this sort of thing (especially a Clojure tuned GC, for I suspect idiomatic Clojure behaves differently than standard imperative Java), but it's a lot of work (needs either a native Clojure implementation, which ought to wait on Clojure-in-Clojure, or working on an existing VM and the fast ones like Hotspot are messy).