I thought the concurrent GC algorithms (zing/azul, g1?) avoided these stop the world events by using concurrency?
With that, STW pause time does not depend on the heap size or on the root set size (the stack size). In practice, it means pause < 1ms, at that point the OS becomes the bottleneck, not the GC.
So latency is good but throughput can be reduced by 30%.