You really don't anymore. For the past several years, Java's GCs mostly pick the right settings automatically, except for heap size, which will be taken care of soon (https://openjdk.org/jeps/8377305). The reason heap size isn't automatic is that with moving collectors it determines the CPU/RAM tradeoff, and doing that in a more natural way isn't trivial, but we have the algorithm now and will merge it soon.
> or making sure to "pick the right collector for the job"
There are really only four options, most of which are easy to choose among: Parallel for batch jobs where only throughput matters, ZGC for interactive applications where latency matters a lot, and then consider either G1 or Serial if there's a problem with those choices.
As someone who's worked for a long, long time solving manual memory management issues, the amount of effort required isn't just in a different ballpark, but in a different city. Sure, spending a few hours a year to reconsider your settings isn't nothing, but it isn't even remotely in the same category of pain with manual memory management (or even automatic memory management, but with malloc/free underneath).
Ten years ago, “rewrite in C++” was definitely easier than getting the Java GC to stay up under server load.
Most servers I work with run on big machines and are the only process, so figure a 100-250GB heap that lives for months, all async, small requests, so insane amounts of Future and String allocation spam.
Optimizing that stuff away in Java is harder than writing Rust, so assume idiomatic Java.
Also, is there any work on statically enforcing data race freedom in Java? That’s a bigger rust selling point than memory safety for me. I think swift has done some interesting work in that space. It would be nice to get those sorts of safety properties without manually writing borrow checker annotations.
Well under 1ms for ZGC (to the point that OS-caused hiccups are of similar magnitudes).
> Ten years ago, “rewrite in C++” was definitely easier than getting the Java GC to stay up under server load.
Both could have been hard in some cases, but open-source "pauseless" GCs are only 3 years old (and all of the JDK's GCs are nothing like what they were ten years ago).
> Optimizing that stuff away in Java is harder than writing Rust, so assume idiomatic Java.
Quite the opposite. Performance issues due to memory management are, in practice, more serious in Rust than they are in modern Java.
> Also, is there any work on statically enforcing data race freedom in Java?
There isn't much demand for that atm. If we see growing demand, we could prioritise it.
The remaining allocator performance problems are mostly due to it zeroing allocated memory unless I use unsafe. Is java able to stackify most new Object calls and elide default initialization of object members these days?
I’m surprised to hear there is no demand for compiler enforced/facilitated thread safety in Java. That was a major pain point in all the Java code bases I’ve worked with in the past, and is a headline safety feature for rust (which goes even further and enforces aliasing rules) and JS. Could you be seeing selection bias in your user base?
You say "just", but this is easy when programs are small. The problem is that this gets harder and harder and harder as programs grow large (the whole point of the JVM's design was to address the performance issues that plague large C++ programs). E.g. someone who works at one of the world's largest tech companies just told me that they have problems with Rust programs spending 30% of their CPU on memory management even when they're as small as a couple hundreds of thousands of LOC.
> Is java able to stackify most new Object calls and elide default initialization of object members these days?
No, the general idea is to just make memory management efficient (although some objects are "stackified" and the compiler will elide zeroing when non-defaults are passed to a constructor). Now, I say "just", but this used to come at the cost of GC pauses and larger footprint. Now it only comes at the cost of a larger footprint.
But there is a definite choice here when it comes to performance. Low level languages give you control that means performance is attained through manual effort. Java takes away control to improve effort-per-performance. Roughly speaking, these tradeoffs mean that when programs are small and the extra effort is manageable, low-level languages are hard to beat, but when programs are large, it is Java that is hard to beat.
> I’m surprised to hear there is no demand for compiler enforced/facilitated thread safety in Java. That was a major pain point in all the Java code bases I’ve worked with in the past, and is a headline safety feature for rust (which goes even further and enforces aliasing rules) and JS.
This used to be a bigger problem when locks were the main mechanism for sharing data among threads. Now, with the wide selection of concurrent data structures, such problems don't occur as much. I'm not saying they don't occur at all, just not frequently enough to become a major priority.
Also, safe Rust's data-race freedom comes at the cost of requiring unsafe for benign races, which are not uncommon in concurrent algorithms (i.e. it excludes even "good" races). This may be fine in languages whose view on performance is "with enough effort you can get good performance", but, as I said, Java is about making more "naive" programs fast with little effort.
Generational moving collectors are a very powerful memory management optimisation, but because they necessarily require some "interesting" FFI layer between the ordinary heap and any passing of pointers between the program and the hardware/OS - the very thing low-level languages are designed not to have - this is a powerful, general optimisation (not perfect, but extremely useful in a wide class of programs) that is not available to low-level programming languages. And it's not the only one, BTW. JIT compilers are also designed for "global" average-case optimisations at the cost of precise low-level control over the worst case, which is another thing that low-level languages trade away.
In the most simplistic way, I would say that the precise control that low-level languages are all about helps their performance when programs are small (and can be manually optimised globally) and hurts their performance as programs get large. They have to give up on some optimisations that come at the cost of ceding low-level control, and that includes moving collectors.
When I say tradeoffs", I mean exactly things like "inherent increased footprint", or my earlier "wasting RAM" point.
At this point, I'm not sure that we're disagreeing about these details, but rather what we make of them... I view them as just years of continued tuning of tradeoffs (heap size vs CPU cycles vs memory cycles), while you them as major breakthrough that makes garbage collection much more desirable. Is that fair?
> When I say tradeoffs", I mean exactly things like "inherent increased footprint", or my earlier "wasting RAM" point.
Well, that is a real tradeoff, but there's a reason why it's a very attractive one for a huge class of applications. There are two ways of looking at this, which amount to the same thing:
1. Because both RAM and CPU are needed for computation, what matters isn't each of their utilisation values separately but only the more constraining or impactful of the two.
2. Because CPU is needed to use RAM, every CPU cycle you spend effectively takes away some other program's ability to use RAM.
This means that for any amount of CPU utilisation, there is some amount of RAM that is effectively free (i.e. has no additional impact), and the more CPU a program consumes, the more RAM it can consume without it making additional impact. It's easiest to see in the extreme case of a program using 100% CPU: no other program can make progress, and so it doesn't matter how much of the available RAM your program is using - it effectively captures all of it whether it uses it or not. But this scales to any amount of CPU utilisation (not quite linearly). What wastes RAM is not using that "free RAM" to reduce the dominant resource, CPU. And what further determines the RAM/CPU "exchange rate" is the RAM/CPU ratio offered by the hardware, which is more RAM-heavy than some appreciate (it is very hard to find a metal or virtual deployment with less than 1 GB of RAM per core these days - taking into account partial cores in virtual machines - outside of embedded devices).
This means that if you have a memory management algorithm that uses more RAM to help reduce CPU as CPU utilisation rises, that's usually a good thing. And moving collectors work exactly like that. The heap overhead in a generational moving collector is only a function of the allocation rate, and a high allocation rate also means high CPU usage.
My colleague, the main developer of ZGC these days, gave a keynote about this very subject at ISMM: https://youtu.be/mLNFVNXbw7I
> I view them as just years of continued tuning of tradeoffs (heap size vs CPU cycles vs memory cycles), while you them as major breakthrough that makes garbage collection much more desirable. Is that fair?
I say that for many years, the main practical, most "felt" tradeoff of moving collectors has been their STW pauses. With pauses eliminated, there is a qualitative change in the attractiveness of moving collectors, making them more appropriate than other memory management techniques for an even broader class of applications than before. Previously, applications that were very sensitive to tail latencies didn't want moving collectors; now, the tail latency is no longer an issue (unless your application's tail latency tolerance is such that a realtime OS is needed). In other words, I'm saying that moving collectors' most impactful tradeoff is now gone.
This is exactly what async programs do on the hot path. Consider a 1M request per second process holding 64K of buffers per request. That’s 64GB of allocations per second. Now, assume the requests hit a remote database with 10ms latency. That’s 640MB of live heap in steady state, which ends up in the “long lived” part of most garbage collectors.
Using RAM to save CPU is exactly the wrong tradeoff when such a system becomes CPU bound.
It’s almost always the case that it is CPU bound due to an incoming request spike or elevated retry rates on the backend. Those tend to pile up, creating a 64GB/sec leak.
The alternative is that the system is CPU bound because the heap is large. This is also very common. Unless each collection takes less work as the heap increases in size, backing off the GC rate to free CPU instantly drives the system into metastable failure, where the GC becomes more expensive because the GC is expensive.
Instead of reasoning about this all the time, it’s much easier (for me, granted, I am not a typical java developer) to just jam the CPU intensive work on a low priority event queue so that it uses 100% CPU but never blocks low latency stuff, or things about to retire requests. (Or, stick it in a dedicated but small thread pool if I can’t touch the async event loops).
This ends up being easier to deal with than java, since everything is thread safe, allocations are predictable, and there are CPU escape hatches I can use.
C++ lets me use smart pointers that have exactly the semantics I want, and that are memory safe but racy in practice. Rust makes them actually memory and thread safe, but sometimes adds useless copies, initializations and thread synchronization (or requires unsafe).
It won't, because that is exactly the thing good moving GCs detect and size the young-gen accordingly.
> Using RAM to save CPU is exactly the wrong tradeoff when such a system becomes CPU bound.
Did you mean to write something else, because it's pretty obvious that it's the right tradeoff? If something is CPU-bound, you want to reduce the CPU usage.
> Unless each collection takes less work as the heap increases in size, backing off the GC rate to free CPU instantly drives the system into metastable failure, where the GC becomes more expensive because the GC is expensive.
The whole point of moving collectors is that each collection takes the same amount of work, but you need to do it less frequently as the heap rises. So yes, as they heap grows, moving GCs are supposed to work less. The heap grows as a function of the allocation rate while the CPU devoted to memory management remains the same. That's precisely the optimisation that moving collectors bring.
> Instead of reasoning about this all the time
The whole point is that the GC is what "reasons" about this for you.
> it’s much easier (for me, granted, I am not a typical java developer) to just jam the CPU intensive work on a low priority event queue so that it uses 100% CPU but never blocks low latency stuff, or things about to retire requests.
That's orthogonal. You can do that at least as easily in Java.
> This ends up being easier to deal with than java, since everything is thread safe, allocations are predictable, and there are CPU escape hatches I can use.
Thread safety is orthogonal, and now with ZGC, memory management in Java is more predictable than malloc/free allocators.
> C++ lets me use smart pointers that have exactly the semantics I want, and that are memory safe but racy in practice. Rust makes them actually memory and thread safe, but sometimes adds useless copies, initializations and thread synchronization (or requires unsafe).
Yes, and it's also less efficient and less predictable in the memory management work as programs grow larger.
Escape analysis in OpenJDK will stack allocate values where it can show it is safe to do so. Project Valhalla is also reducing the memory footprint of objects.
As for thread safety, that is more of a language concern than a runtime one. Amongst JVM languages Scala is leading here AFAIK. Its "capture checking"[1] provides thread safety (e.g. [2]) and actually covers escape analysis as well. On Scala Native (the native code backend for Scala) capture checking can be used for safe stack allocation and safe arena allocation.
[1]: https://docs.scala-lang.org/scala3/reference/experimental/cc... [2]: https://softwaremill.com/understanding-capture-checking-in-s...
My reason for disagreeing with the former view is that improvements in physical RAM available and tendency towards smaller workloads have allowed many Java (or other GC runtimes) to essentially "fix their problems because hardware got better". So you can waste more RAM, waste more cycles, but "it doesn't matter", and likely it is fine in many cases - but it's is not the same thing as claiming the GC algorithms are responsivle for that outcome. We have been 3 years away from GC solving memory management for at least 30 years.
My reason for disagreeing with the latter view is that for those who don't have 100-250 GB long-lived heaps (or whatever the contemporary version of that is), the pain level is far lower than rewriting in C++ or Rust, or likely the pain level of hiring enough engineers who can do either. It's a completely different engineering culture.
I’ve also worked on systems with lots of small processes, and the operational issues that creates dwarfs GC problems: It takes one middle tier machine, and adds 64-128 network boundaries, and also creates an extremely difficult static memory allocation problem.
I know people do it anyway, but it’s rare that they can articulate a decent technical reason for it, and it wastes something like 90% of the hardware (even in carefully optimized code bases / deployments).
Anyway, I’m not the target market for such stuff.
Just to be clear, the main reason for the use of moving collectors in the first place is to waste less cycles on memory management (otherwise we wouldn't use them). They exist to serve as an optimisation.
> We have been 3 years away from GC solving memory management for at least 30 years.
It's now 3 years in the past (since Generational ZGC); e.g. see https://netflixtechblog.com/bending-pause-times-to-your-will.... Of course, it doesn't solve all imaginable memory management issues, but in practice it makes it a non-issue for a large class of interesting and very common programs.
Even in the very positive blog you linked, you see statements like * "ZGC has a fixed overhead 3% of the heap size, requiring more native memory than G1. .." and * "Reference processing is also only performed in major collections with ZGC. We paid particular attention to deallocation of direct byte buffers, but we haven’t seen any impact thus far. This difference in reference processing did cause a performance problem with JSON thread dump support, but that’s a unusual situation caused by a framework accidentally creating an unused ExecutorService instance for every request."
This was my point about how this sort of thing is a type of manual memory management.
As for waste more RAM, waste more cycles wasn't a statemnt about whether a particular GC is better-performing for certain situations, but that the overall improvement likely has more to do with improvements in CPU speeds and RAM size, than the latest GC version (which tends to simply make a different set of engineering tradeoffs).
> but that the overall improvement likely has more to do with improvements in CPU speeds and RAM size, than the latest GC version (which tends to simply make a different set of engineering tradeoffs).
Well, the biggest improvement has been the creation of a new "pauseless" collector, ZGC, with a novel GC algorithm (at least for OpenJDK), which does _zero_ GC work in STW pauses, i.e. no scanning, no marking, no moving. In particular, even roots, including stacks, are processed entirely concurrently with the program. The main practical impact of that has been saying goodbye to GC pauses, and getting low latency, that is perhaps even more predictable than malloc/free (and obviously, still has higher throughputs in a large class of interesting programs). The tradeoff is the usual footprint tradeoff, which is the core of moving algorithms, as well as more CPU cycles compared to STW collectors (but again, still less than malloc/free in many programs). The additional CPU can, of course, be compensated for with an even larger heap.
The general idea is to use RAM chips as hardware program accelerators, but in the past latency was also something you had to sacrifice, and this is no longer the case today.