So the actual tradeoff is more whether you want better throughput or tail-latency. To improve the latter, you have a singular command line option of using ZGC as well.
[1] https://inside.java/2022/05/30/sip053/#:~:text=ZGC%20was%20d....
[1] Reduced them to always well below 1 millisecond.
This can be significant too.
Java's GCs are incredible and you have a menu of algorithmic options that let you avoid whatever problem you're worried about.
Today we've removed CMS, added G1GC, and added ZGC.
G1GC is similar to the parallel in the way it operates but is divided up enough to control for latency.
ZGC is fundamentally different from the parallel collector or G1GC. Suggesting it is the same is to suggest you know nothing of the changes that have happened.
can these options be used in separate parts of a single app ? e.g. the app I'm developing in C++ has parts that do realtime audio, others that do GPU rendering, others that do classic Qt Widgets GUI, others that do offline computations on datasets - and they all have different performance characteristics and need different memory management schemes to get the best out of each, all while being in a single process; there's reference counting, tree-based allocation, pooled, linear, a GC-ish thing which ensure that memory is freed in specific non-realtime threads... Can that be done with Java or is one tied to a single GC implementation for a given execution of a process?
- In the "in-process" case I can stack ~1400 plug-ins on a single channel before I hear a crack in the sound.
- In the "shared memory" case I can go up to ~200 at most. And I'm confident that they really did the very best things possible for the implementation to be performant.
So for me the "things isolated in their own process" means literally getting seven times less out of my system than in the host process (and that's frankly unuseable, definitely not "negligent").
1. Instant startup due to AOT compilation and a cached heap. Can start faster than C!
2. No warmup.
3. Can create native code shared libraries.
4. Offers isolates, which are segregated heaps that do GC separately but run in-process and which can communicate with each other.
The tradeoffs are that unless you buy the more advanced edition, peak performance is lower due to lack of JIT profiling, you may need to write configs and do other fiddling to ensure the AOT compilation doesn't miss any code that's accessed via reflection, it takes a long time to compile, and you can't dynamically load bytecode (which some libraries do behind the scenes transparently).
True concurrent GC have been available since late 2000s. Azul had a read-barrier GC as well - effectively a pauseless GC (or pauses under 1ms)
This is a feature, not a bug. Which is riskier?
While the JVM offers a plethora of levers to pull in case you are hyper concerned about different things, the heuristics are VERY good. Mucking with the fine grained details can disable heuristics and ultimately give you worse results than if you just left stuff alone.
The levers to pull are algorithm, max memory, and max pause time. All other levers should be left alone unless you've got GC logs to back up what you think needs changing. (And even then... Do you really?)
Typically, the better route is flight recorder and eliminating wasted allocations.
Lisp I Programmers Manual, 1960
https://www.softwarepreservation.org/projects/LISP/book/LISP...
Chapter 6.3, Page 95, The Free-Storage List and the Garbage Collector
Fwiw, the JVM now has a noop garbage collector so this is easy enough to benchmark.
But it is easy to monitor with Java’s killer observability.