ZGC – What's new in JDK 16
malloc.se
malloc.se
Even if the GC is running on an otherwise idle core there are still other costs like power consumption and memory bandwidth. So you still want to minimize allocation to keep the GC workload down.
For too long GC people were touting 10 ms pause times as "low" and not bothering to go further, but truly low pause times are possible. I'd love to see a new systems language that starts by designing for extremely low-pause GC, not manual allocation or a borrow checker. I think it would be possible to make something that you could use for real time work without having to compromise on memory safety and without having to pay the complexity tax Rust takes on for the borrow checker.
Though you can configure for this in a couple ways if you run into this issue
1. By telling it to treat some amount of heap as the max that is not actual max when it comes to it calculating when it should start next gc cycle. (-XX:SoftMaxHeapSize) https://malloc.se/blog/zgc-softmaxheapsize
2. Increasing the amount of concurrent gc threads so they will finish their work faster (-XX:ConcGCThreads)
3. Just run really large heap so the "run a gc cycle even if we don't need to based on allocation rate" gc cycle keeps the heap in check.
Though after JDK 15 we have not had to mess with any of these. Prior to that we had to adjust soft max heap size a bit. With JDK 16 it should be even better I guess (should be upgrading sometime next week)
A single skipped frame is not a big deal and will probably not be noticed. It will probably happen anyway due to scheduling quirks, resource contention with other processes, existing variation in frametime...
True realtime work requires no dynamic allocations whatsoever (which, notably, is not covariant with gc!), so I think ‘low’ pause times are an acceptable compromise. Where performance is a concern, you need to manually manage a lot of factors, among them GC/dynamic memory use. There's no runtime that can obviate that.
Granted, 1ms pause times are probably still not low enough for realtime audio, and there may be room for some carelessness there (audio being soft realtime, not hard realtime). But I think just being careful to avoid dynamic allocation on the audio thread is probably a worthwhile tradeoff.
Attitudes like this are why my phone sucks to use and I get nauseous in VR and GC devs spent so long in denial saying 10 ms pause times should be good enough. Yes, single dropped frames matter. If you don't think so then I don't want to use your software.
It won’t make the normal case jittery, nauseus or anything like that. Also, in regards to your GC devs comment, I would say that attitudes like this is the problem.. The great majority of programs can do with much more than 10 ms pause times.
It's in the same category as "the performance of stock Camry is unacceptable for racing" - yes, plan accordingly when entering a race :-)
(Though there was an allocator I saw recently that promised O(1) allocations. Pretty neat idea.)
A bump allocator that you reset every frame is O(1) and a dozen or so cycles per allocation for example.
Some folks definitely notice this phenomenon, called a "microstutter" by that group. You can see it here:
https://testufo.com/stutter#demo=microstuttering&foreground=...
Is a single frameskip in an hour a problem?
Unless you have a highly optimized game, you are probably not able to consistently run at a 144 Hz monitor’s native refresh rate anyway, so even without skipping frames you will see stuttering. VRR solves this problem as well.
Of course, we’re generally talking about much higher refresh rates, at least 120 if not more.
My display runs at 165hx and you’d be surprised how many games could hit that a lot of the time... even on my old 1080 (and certainly on the 3080 I have now).
The most frequent issue I run in to is actually games having hard FPS caps (120 typically, sometimes 150) and refusing to actually max it out.
1ms (max, average at 50us)
And for the 'when' I'll add that the very concept of having a concurrent GC means you don't need to do a (potentially pausing) malloc right in the middle of what you're trying to do.
Unreal is more common than it was 5 years ago, but it's still a very distant 2nd. Probably 9/10 games released on Steam in 2021 are Unity.
You may not realize it, since the devs are slightly less bad at hiding the fact thay they're using Unity. (How did anyone at Unity ever think that disaster of a "launcher" was a good idea?)
https://docs.unrealengine.com/en-US/ProgrammingAndScripting/...
"Garbage collection in Unreal Engine 4 is fast and efficient, and has a number of built-in features designed to minimize overhead, such as multithreaded reachability analysis to identify orphaned Objects, and unhashing code optimized to remove Actors from containers as quickly as possible."
I always feel like there are a LOT of posers in every HN/reddit discussion of GC who have tricked out gaming rigs and are desperate to max it out in every game they play, but aren't actually game developers themselves. Whilst out in the real world there is Unreal, Unity, heck even Minecraft all making mad coin using garbage collectors at their core.
Coming from the demoscene and having had a foot on game development, I am fully aware that stuff like monetizing IP and getting a game published no matter what and how, are much more relevant in the game development communities than whatever gets discussed over here.
Just look at e.g. the general opinion about Flash around here, and how it was embraced by the game development comunities.
On Reddit there are gamedevelopment forums where you will have more luck finding more real life experience.
Then there is IGDA, Gamedev, Gamasutra, Making Games, PAX, IGF and similar forums.
Fact is, not every game engine needs to be used for a Far Cry clone, and there are many ways to make money.
It was the worst GC implementation I've seen in my life, could cause 0.5s GC spikes every 10 seconds on Xbox One even though we were allocating none or very little memory during gameplay. The amount of pre-allocated and pooled objects was bringing it down to its knees, because Unity's GC is non-generational and checks every single object every time. In the end we moved a lot of the data into native plugins written in C++. Nothing super hard, but you choose high-level engine to avoid such issues.
I've read that in 2019 they finally added incremental mode GC, that solves some of the issues but is still far cry from modern GC's.
1. Try to bump a pointer to allocate the memory
2. If there wasn’t enough memory, call a special function (which runs the GC) and try again
This special function may do a big gc pause if necessary or it may just do a small amount of gc work and grant the program some more memory (this way most gc work is shared more evenly with program time). In modern VMs like Java’s, concurrent gc (where gc and mutator run in separate threads possibly needing a pause for some steps) is more popular.
In a game an object is likely to live for exactly one frame, exactly n frames where n is like 2 or 3 for some kind of deferred rendering, or an entire level. Some games mightn’t have levels as such but objects still tend to be grouped together in when they load/unload.
I think it is important to note this because games that need to care about performance are not using malloc. They will use some kind of arena allocation because allocating and freeing memory with malloc can be expensive. So what games want here is control over how allocation happens rather than when free is called—that is, it wouldn’t be sufficient to have a language where you called free but couldn’t use a custom allocator.
But there are standard ways to deal with these sorts of allocation patterns in a language like Java. For buffers (of eg vertices) these may just be reused. For smaller objects, pool allocation may be used which sucks but doesn’t suck much more than manual memory management.
I think it is silly to focus on games (or, for that matter, short-running command-line tools that needn’t call free at all) because the allocation pattern is so unusual and weird. I can’t think of great examples of systems that care about latency and don’t have game-like allocation patterns but maybe a complex gui application like a web browser or PowerPoint would be a good example.
This means you may not activate the compressed pointers optimization.
Ignoring pause times is fine for batch processing, but not ideal for interactive systems.
This is a fundamental principle of garbage collection. You can either have low latency or high throughput. You can't get both.
Why is that?
All optimizations that improve latency come at a cost. Generally, more book keeping, more checks, more frequent garbage collections. ZGC is one of those algorithms. It adds a new check every time you access memory to see if it needs to be relocated. That increases the size of objects but also the general runtime of the application.
A similar thing happens with reference counting (which is on the extreme end of the latency/throughput tradeoff). Every time you give a shared pointer or release a shared pointer a check is performed to see if a final release needs to happen.
On the flip side, a naive mark and sweep algorithm is trivially parallelizable. The number of times you check if memory is still in use is bound by when a collection happens. In an ideal state you increase heap size until you get the desired throughput.
We get "violations" of some of these principles if we can take shortcuts or have assumptions about how memory is used and allocated. For example, the assumption that "most allocations are short lived" or the generational hypotheses leads to shorter pause times even when optimizing for throughput without a lot of extra cost. It's only costly when you've got an application that doesn't fit into that hypotheses (which is rare).
Haskell has a somewhat unique garbage collector based on the fact that all data is immutable. They can take shortcuts because older references can't refer to newer references.
When you don't mutate in OpenJDK you get essentially the same. Much of the cost of a modern GC (OpenJDK's G1, and soon probably ZGC, too) is write barriers, that need to inform the GC about reference mutations. If you don't mutate, you don't pay that cost. This is partly why applications that go to the extreme in the effort not to allocate and end up mutating more, might actually do worse than if they'd allocated more with OpenJDK's newer GCs.
In fact, OpenJDK's GCs rely heavily on the assumption that old objects can't reference newer ones unless explicitly mutated, and so require those barriers only in old regions.
If you go to the extreme and don't allocate, you can turn the GC off.
I don't think this is true? Because laziness is really heavily mutable under the hood. Not to mention that it has mutable references. But maybe there are some tricks in the GC I'm not aware of.
> All optimizations that improve latency come at a cost
Huh? That's a very strange, absolutist statement. There are many cases where you have the opportunity to trade off one of latency and throughput for the other, yes. But there are also many cases where an optimization can improve both.
AFAIK, there's no algorithm that's both good at throughput and latency without making assumptions around how memory is used.
For many systems this is desirable. It’s also the reason that incremental and concurrent gc were invented even though you could just have a gc where the mutator runs until it can’t allocate anymore, then you do a big mark and sweep, then you resume the mutator.
Decreasing tail latency can generally improve downstream systems. Let’s say latency for requests is iid and I have some intermediate system which receives a request, needs to send 20 requests to your server, wait for them all to complete, then responds. Then the median latency of my system will look like the 95th %ile latency of your system that it gets data from. If your system were a web server and immune a web browser then you could pay yourself on the back for having low median latency while your users actually experience your 95th %ile latency waiting for all the resources to load.
This is nuts (and very well below OS jittering)
Bear in mind that by default most Linux distros will a timeslice of 10 milliseconds so your thread can easily skip that much time simply because it was some other thread's turn to run. JVM pauses with ZGC and even G1 most of the time (when set to the most aggressive settings) are now so low, that your pause times will be dominated by other factors.
Say I'm doing a drawing/game app and creating a few hundred heap objects a second that need to get garbage collected.
I have no idea on how often GC is run on a typical app and how much real time it takes over say an hour of an semi complex app running on average. It obviously depends on the app but I do not even have a number average cost of a GC language for some typical web app.
I only know 'GC's are bad' because the 100s of HackerNews comments dismissing languages because they have a GC for some reason rather than hard examples of them eating up time.
In past GC had bad reputation for increased and unpredictable latencies. In old JVMs GC would pause execution to traverse object graph.
In general do not worry about GC, unless you run into performance issues. If performance is a problem, run continuous profiler such as Flight Recorded. It has very little overhead.
GC has many different implementations, with widely ranging properties. For example, the JVM itself currently supports at least 3 different GC implementations. There are also different types of GC's, so for example in a generational garbage collection system you'll typically see two or three generations of GCs, depending on the generation (how many GC cycles it has survived) of the objects it collects. The shortest GC's in those systems are usually a couple milliseconds, while the longest ones can be many seconds.
GC isn't always a problem. If your application isn't latency sensitive, it's not a big deal. Though if you tune your network timeouts to be too low, even something that is not really latency sensitive can have trouble because of GC causing network connections to timeout. Even if it is a latency sensitive applicatoin, if GC "stop the world" pauses - pauses that stop program execution, are short it can be OK.
One reason you'll see people say GCs are bad is for those latency sensitive applications. For example, I previously worked on distributed datastores where low latency responses were critical. If our 99th percentile response times jumped over say 250ms, that would result in customers calling our support line in massive numbers. These datastores ran on the JVM, where at the time G1GC was the state of the art low-latency GC. If the systems were overloaded or had badly tuned GC parameters, GC times could easily spike into the seconds range.
Other considerations are GC throughput and CPU usage. GC systems can use a lot of CPU. That's often the tradeoff you'll see for these low-latency GC implementations. GC's also can put a cap on memory throughput. How much memory can the GC implementation examine with how much CPU usage with what amount of stop-the-world time tends to be the nature of the question.
But the main thing most folks have problems with is the random latency spikes you get with GC. The GC can start at any time in most languages, and might stop all threads in your program for maybe dozens or hundreds of ms. This would be visible to users if you are rendering frames at a constant rate in a game, since each frame takes only around 16 ms in a 60 FPS game.
That’s what’s exciting about changes like what they are doing with ZGC. They are saying the max garbage collection time is 0.5 ms in normal situations, and the average time is even lower. Most games can accommodate that without a problem.
FYI, this is also important for web servers as well. Some web servers have a huge amount stored in memory, and the GC could take hundreds of ms or even multiple seconds to collect at random times in extreme cases. This can make a web request take perceptibly longer.
Also, if you have multiple machines communicating with one another and randomly spiking in latency due to GC, then worst case latency can add up to pretty terrible numbers if you are not careful.
The thing about GC is you either don't care at all, or you don't want it at all. There's rarely a case where you know how many GC cycles you can handle in a certain period. Web dev, GC all you want. Games can handle GC but its likely you'll need to be cognitive of memory use. Embedded stuff doesn't have enough memory to utilize a GC.
I'm sure why GC languages get so much hate. I do a lot with C# and the runtime gives a few options for controlling allocations and accessing memory, so I can usually get it to be fast enough.
Was literally my job ten years ago to optimize this and I was struggling with a GC'd language with a proprietary implementation (flash+actionscript).
The problem is not with hundreds of heap objects per-frame, the problem is that they would accumulate to the tens of thousands before the first GC trigger happens.
And the GC trigger might happen in the middle of drawing a frame, even worse, at the end of drawing a frame (which means even a 10ms pause means you miss the 16ms frame window at 60fps).
The problem that most people had was that this was unevenly distributed and janky to put it in the lingo. So you'd get 900 frames with no issues and a single frame that freezes.
So most of the problem people have with GC pauses is the unpredictability of it and the massive variations in the 99th percentile latency in the system, making it look slower than it actually is.
Most of the original GC implementations scale poorly as the memory sizes went up and the amount of possible garbage went up, until the GC models started switching over the garbage-first optimizations, thread-local alloc buffers and survivor generation + heap reserves etc (i.e we have lots of memory, our problem is with the object walking overheads - so small objects with lots of references is bad).
The GC model is actually pretty okay, but it is still unpredictable enough that tuning the GC or building an application on top of a GC'd language which has strict latency requirements is hard.
However, as a counterpoint - OpenHFT.
Clearly it is possible, but it takes a lot of alignment across all the system layers, but at that point you might as well write C++ because it is not portable enough to run anywhere.
I recommend this paper from 1992 for an introduction of what types of GC techniques exist https://www.cs.rice.edu/~javaplt/411/15-spring/Readings/wils... It's good basic knowledge.
If you get through that you might want to peruse more modern techniques: https://gchandbook.org/ I think the book's a lot drier/harder to understand than the paper though.
As a counter-signal consider my comment to be dismissive of languages without a GC, because life is too short to deal with manual memory management outside of where you really need to do it, and in those cases anyway there are techniques you can use in GC languages for it that reduce to about the same effort as manual memory management.
Sorry for my ignorance on the topic, but will this have any impact on other JVM languages or will this mostly only benefit Java itself?
I realize even though I use JVM languages now and then I do not really know if they use their own GC implementation or make use of Java's. Does this differ between the languages maybe?
There may be tiny differences in the way code generators and optimizers work, which mean they may not get exactly the same properties out of equivalent code. For example, if they're generating a lot of objects behind the scenes, the GC improvements might help more, or less, or even do worse.
But that's the kind of thing that's really dependent on the algorithm you've implemented. So mostly likely you get some benefit for free. If you don't, you'll need to benchmark to find out. The optimizers do a lot of work for you (and the JVM does a ton of language-independent optimization), but some things are up to experiment.
https://jet-start.sh/blog/2020/06/09/jdk-gc-benchmarks-part1
I'm curious about this choice. The elasticsearch documentation recommends a maximum heap slightly below 32GB [1].
Is this not a problem anymore with G1GC/ZGC, or are you simply "biting the bullet" and using 92G of heap because you can't afford to scale horizontally?
1: https://www.elastic.co/guide/en/elasticsearch/reference/7.11...
Doesn't have to be because of affordance but rather it's more efficient and cheaper to scale vertically first, both in monetary costs and in time/maintenance costs.
In cloud, with defined instance types, usually more ram comes with more everything else, and from pricing listed at https://www.awsprices.com/ in US East, it looks like within an instance type, $ / ram is usually consistent. The least expensive (per unit ram) class of instances is x1/x1e which are 122 Gb to 3904, so that does lean towards bigger instances being cost effective.
Exceptions I saw are c1.xlarge is less expensive than c1.medium, c4.xlarge is less than other c4 types and c4 is more expensive than others, m1.medium < m1.large == m1.xlarge < m1.small, m3.medium is more expensive than other m3, p2.16xlarge is more expensive than other p2, t2.small is less expensive than other t2. Many of these differences are a tenth of a penny per hour though.
From my experience, high heap sizes are unnecessary since Lucene (used by ES) has greatly reduced heap usage by moving things off-heap[1].
[1] - https://www.elastic.co/blog/significantly-decrease-your-elas...
AOT compiled, PGO optimized, statically linked executables, with 0.5ms worst case GC pause times, sounds like the holy grail for me.
Note that Oracle does have intentions to commercialize similar features; there's already a low latency GC for native-image apps, but it's inspired by G1, not ZGC, I believe, and only available in Graal Enterprise Edition. So any such ultra-low-latency GC might be in a similar basket, unless it was implemented by an outside party or something. Or Oracle changes course on this.
If you don't use native-image and can use a normal JDK to host your Graal app (by using Graal as the JVMCI compiler), then you might be able to use ZGC today with that setup, assuming support for ZGC read barriers has been implemented in Graal.
EDIT: apparently not yet :( https://github.com/oracle/graal/issues/2149
-XX:+UseZGC: Memory usage for my project dropped to a constant 600 megs. Using the IDE felt just as fast as the normal experience.
-XX:+UseShenandoahGC, -Xmx4g. Shenandoah GC used a constant 4 gigs of ram. It was a slower user experience for me.
In the end, I went back to the default settings, because the custom JDK changes the look and feel and I don't like it.
Shame, I hoped it would feel faster than the normal experience, with (even infrequent) user-felt GC pauses completely eliminated.
Very impressive and well done. Should Azul be worried?
https://github.com/openjdk/jdk/tree/master/src/hotspot/share...