Azul launches pauseless gc jvm for unmodified linux
azulsystems.com
azulsystems.com
It's disconcerting that they need to mention "512GB", as if they found that "pauseless garbage collection was" not a big enough selling point (perhaps because for most server applications - where the big money Java apps are - individual responsiveness doesn't matter all that much).
It's also disconcerting, because (it seems to me) that big apps are increasingly a thing of the past, as we move to massively distributed smaller, slower processes - each individual JVM with its own smaller memory needs. Big apps seem to be an upmarket niche, disappearing fast. Whereas the great strength of pauseless gc is in realtime client-side apps - like mobile devices.
I think these guys have done something great here, and it would be a shame to see them pushed upmarket into oblivion, when they could create an entirely new disruption based on Java - perhaps (finally) fulfilling its promise... (and for the java-haters, also for other JVM-based languages).
EDIT Azul tech previously on HN (links include explanations IIRC): http://news.ycombinator.com/item?id=810506 --- http://news.ycombinator.com/item?id=2022723 --- http://news.ycombinator.com/item?id=2058476
Also the other thing to remember is that the standard java GCs dont reliably work under all workloads above 2-4GB. You can make a 8gb or 16gb heap work, but if you hit a full-compaction your JVM will go MIA for MINUTES at a time. So you can't even scale to a 16gb ram heap.
The 512GB is for the previous customers who do in fact use heaps of that size.
http://www.azulsystems.com/products/zing/c4-java-garbage-col...
The Azul collector doesn't need the changes in page protection to be visible immediately, so their kernel patch allows batching together the changes and committing in a single operation.
As far as I can tell from reading their whitepaper, however, they still require kernel patches.
a) Ship the source code alongside the binary
b) Offer to furnish the source code to people who ask for it
AFAIK, the second option was intentionally introduced to provide for a delay in releasing source code, since providing the source code to uninterested parties might be cost-prohibitive if it's 1992 and you're publishing on floppies.
But just one thing... according to Wikipedia [1] it is actually based upon Hotspot, but they also entered an undisclosed agreement. In such aspect, Azul may have licensed Hotspot not through GPLv2 but through different terms with Oracle. If that is true, they may be able to keep part of their code out of GPLv2.
Also they had some open source stuff released under GPLv2 at another website [2].
[1]: http://en.wikipedia.org/w/index.php?title=Azul_Systems&o...
Edit: I upvoted you moonboots, thanks for the pointer below.
- Cliff Click, http://www.infoq.com/interviews/click-gc-azul
Great, so you can patent a GC algorithm now.
Instead of deferring the hard cases as long as you can (e.g. in generational GCs doing collections beyond the nursery) at which point they're likely to be painful, concentrate on the hard cases first and foremost. Get those right and everything else falls into place. Here's their CTO discussing that in a short interview: http://www.artima.com/lejava/articles/azul_pauseless_gc.html
A relevant quote:
[...] Our collector really does the only hard thing in garbage collection, but it does it all the time. It compacts the heap all the time and moves objects all the time, but it does it concurrently without stopping the application. That's the unique trick in it, I'd say, a trick that current commercial collectors in Java SE just don't do.
Pretty much every collector out there today will take the approach of trying to find all the efficient things to do without moving objects around, and delaying the moving of objects around—or at least the old objects around—as much as possible. If you eventually end up having to move the objects around because you've fragmented the heap and you have to compact memory, then you pause to do that. That's the big, bad pause everybody sees when you see a full GC pause even on a mostly concurrent collector. They're mostly concurrent because eventually they have to compact the heap. It's unavoidable.
If they do does anybody know more about whats going on in terms of moving these features into the kernel? I could really find anything on it.