Maximum STW of <1ms? Uhh, ok. The best-of-breed open Java GCs have targets of <100ms on big heaps, 1ms ceiling sounded insane.
However, diving into it, the problems solved by the Go GC and those solved by the Java GC are very different in at least two major ways:
- Go has substantially less dynamic runtime behavior; there's no JIT, no runtime loading of new code, very limited reflection, &c.
- The Go GC does not do heap compaction
Both of these things mean: a) Objects do not randomly move about the heap and b) code does not randomly redefine itself.
Which, the Go team have made their bed, just like the Java team have, and the Go bedroom looks a heck of a lot more comfy to sleep in. Building a fast GC for Go is a massively simpler problem than building a GC for Java, hence the claims they are making of worst-case scenarios are not completely off the wall.
Now, that doesn't mean they should be taken at face value. Notably, the Go GC still has to decide when to start a cycle, and still needs to stop the application if it runs out of heap due to starting a cycle too late. I would suspect that you can construe a test case with a very large heap and a random and extremely varying allocation rate that could cause the Go runtime to explode just fine.
Likewise, the time to complete a GC cycle is (at best) linear to the heap size of Go - as the heap grows, the time to reclaim memory does as well, meaning you need to keep larger margins in the heap for floating garbage waiting to be reclaimed.
I was writing the Java part of a cross-platform game engine once, and I was trying to eliminate all dynamic allocations during the game (front-loading all allocations and then never touching the heap). Did some profiling and determined I'd missed something.
The allocation turned out to be in a call to get the current system timer (in milliseconds, if I recall). I was calling a Java function that returned a Java long int, probably straight from an OS call that returned a 64-bit int, and somewhere in that call it was allocating an Object.
Ripped out the Java library call, replaced it with a JNI call to a two line C function that returned the result of the corresponding OS call. No more dynamic allocations.
It's just part of the philosophy of Java to allocation things everywhere. So it's not only hard to make a GC fast for architectural reasons; the GC also tends to be worked harder by the way Java code tends to be written.
You're also hitting another nerve: The horrific world of allocation tracking in Java. While your case sounds like a legitimate thing, so many times I've been in that situation and been tricked by my tools to optimize the wrong path.
Most popular allocation tracking tools, like YourKit, turn off escape analysis. Obviously, when you do that, stupid things like boxed primitives will show up as your main sources of heap allocation even though Hotspot would stick those suckers on the stack in a heart beat.
Eg: Here's a program running without YourKit allocation tracking: https://twitter.com/jakewins/status/707759792547233792
And here's the same program with it on: https://twitter.com/jakewins/status/707754868497231872
Still makes me furious - what in the world am I supposed to do with an allocation tracking tool that creates millions of false positive allocations, drowning out the thing I care about? Gah.
I like the language, but it did a bit of left turn by not adopting the features of Oberon(-2), Modula-3, Component Pascal or Eiffel in regards to value types and native AOT compilation.
https://www.azul.com/press_release/azul-systems-extends-supp...
And with Shenandoah [1] too!
Also, everyone talks about heap sizes; which is kind of the size of the problem if you let it accumulate. I would be much more interested in having these latencies measured in the context of sustained allocation throughput, with varying object lifetime lengths. How much the system can take while staying afloat? Here is the most performant gc: never allow to allocate anything.
Go programming constrains heavily towards message passing. I guess this is some kind of tradeoff. I really hope it still makes a lot of new apps possible (40k users concurrently interacting multiplayer games, anyone?).
For example Wooga has quite a few games using Erlang, Hallo used C#/Project Orleans, just to cite two examples.
I meant not in an embarrassingly parallel worload. The 40k would be with interacting with each other in realtime, and not by groups maxed at 10-20. Each of the 40k being impacted by the other 40k's actions; with an all-reduce step between them at each tick. Each of them getting a different overview of the other 40k.
Imagine a FPS where users receive updates on others' position 60 times per second if they are close, 1 time per second if they are far away, and none if the player is not looking at them. This could enable vastly larger games.
The question at hand would be: Go and Erlang's GC enable great latencies; but can their programming model still enable this worload easily? What kind of degree of shared mutable state do we need?
[1] http://highscalability.com/blog/2014/2/26/the-whatsapp-archi...
https://blog.pusher.com/golangs-real-time-gc-in-theory-and-p...
where G1 still had some reasonably large pauses FWIW. Which is surprising isn't its billing to not do that? But FWIW.