JDK 20 G1/Parallel/Serial GC Changes
tschatzl.github.io
tschatzl.github.io
ORCA (and Pony language) solved this while allowing selective mutability and zero-copy message passing improving on Erlang HiPE/BEAM by tying objects to a tiny heap with individual actors (cooperative async threads). There is no global locking in Pony except in limited circumstances.
Zulu C4 is an improvement. Schism and Metronome are less pausey but slower overall.
https://www.azul.com/products/components/azul-zulu-prime-bui...
https://dl.acm.org/doi/10.1145/1809028.1806615
https://researcher.ibm.com/researcher/view_group_subpage.php...
This is why I don't understand the WASM GC proposal, which I understand to be an attempt to make a GC that works for all languages. Can you really write a GC that performantly supports both Java and Go given the different tradeoffs/approaches each makes with respect to memory management, layout semantics, etc?
I mean, the Graal JVM GC supports a ton of languages, especially if you include their LLVM support. Many of them are quite high performance. Truffle Ruby I believe is still (one of if not) the fastest Ruby implementation(s).
It makes Java's GCs quite Java-specific in practice because there aren't that languages that even moderately widely used which are both statically typed (to the degree that it's possible to tell pointers apart before generating code) and restrict pointers to point to the start of the object.
Also supporting interior pointers is not too cumbersome even with a "Java-style GC" (whatever that exactly means). It requires an additional bit per smallest aligned object size to denote if an object is allocated or not. If you need to check if a pointer is an interior pointer to an object, you walk back in the bitmap and find the last object with a bit set to 1 (i.e. find the allocated object which was right before the pointer you have), get the size of that object and check that the pointer you have is within the range of the start and end of the object.
EDIT: The big issue you would face with directly porting a Java GC into other languages is the weak reference processing semantics. Each language has their own subtle differences which would make it a pain.
Keep in mind that typical builds of Hotspot support four or five garbage collectors, and most of them with two different pointer sizes. The required read and write barriers differ widely between collectors and in some cases even collector modes.
I'm not saying that it's going to be easy (an incredible amount of effort went into Hotspot over the years). But it's definitely possible to support wildly different GC strategies efficiently with sufficiently late code generation.
Webassembly, as I see it, is about creating VM which provides acceptable performance with ultimate sandboxing.
It's a replacement for JavaScript transpilation, not for JVM.
Of course it's tempting to do it all without sacrificing anything, but I guess you can't have a cake and eat it too.
[1]: https://users.cecs.anu.edu.au/~steveb/pubs/papers/g1-vee-202...
That is the cool thing about having multiple implementations.
I heard not so good things about Herb Schildt's book.
The argument against exceptions is that non-local control flow can introduce obscure bugs, like forgetting to clean up resources on the (invisible) exceptional execution path. On the other hand, without exceptions, it's possible to end up with no error handling at all by accident, and that too is not visible in the source code (but linters can help, of course).
Many languages which are anti-exception as a matter of principle still use them to report out-of-memory conditions, out-of-bounds array indices, integer division by zero, or attempts to access absent optional values. Not doing this results in code that is simply too verbose.
Except this is still possible on the non-exceptional execution path. You simply just need to forget the defer call. The only thing that solves this is RAII and destructors.
Probably because you wrote, "At my workplace, we’re thinking about using it" and not "my workplace is demanding we use it"
> or that I haven’t tried several others
(I'm not the person who replied and suggested Go, but your comment does make it sound like you are in the language-evaluation phase, not the "we're fully committed to Java after reviewing multiple options" phase)
"Is there a good guide to learning modern Java?"
The end. YOU obfuscated your own comment with everything after it.
And no, the vast majority of use cases for java isn't distributing binaries or interoping with JNI.
Java in 2023 is not what you may remember.
1. The Java Programming Language book:
The K&R of Java
https://www.amazon.com/Java-Programming-Language-4th/dp/0321...
2. Zantorc's videos: https://www.youtube.com/@Zantorc/
He really gets into the nitty gritty. You can listen to this while commuting, etc; Just finish 1 video per 3 days, say. It adds up quickly.
When is JetBrains coming out with KVM?
Guest language everywhere besides Android.
For low resource use cases, you probably want Graal AOT above all else... which means G1GC or serial, I think.
>With 1 core it is always serialgc
Even with 1 CPU ParallelGC has lower latencies than SerialGC on 1 CPU. SerialGC will be better on environment with limitations on number of threads, not number of CPUs.
> 2 cores and less than 4 gb - concurrent mark and sweep IIRC
CMS has been deprecated in 11 and removed later in a non-LTS release. JVM ergonomic will automatically turn on >G1GC< when it detects JVM has at least 2 CPUs and 1792Mb of RAM (not heap, memory in total). When either or both numbers are lower then ParallelGC is enabled automatically.
[0] https://docs.oracle.com/en/java/javase/18/gctuning/introduct...
Otherwise, start with G1 and get your Xmx value in the ballpark. VisualVm can help you determine if you're thrashing. Are you GCing like 10+ times per second? Keep an eye on it. If you start hitting giant pause times and 50+ collections a second, you've got problems :) Increase Xmx. (and no, please don't set Xms = Xmx).
If you have issues past that, it’s not the garbage collector that needs help; the next step is to audit your code base. Bad code makes any Garage Collector in any language to misbehave.
For instance, are you `select *`ing from a table then using the Java streams api to filter rows back from a database? That will cause GC issues :) fix that first.
So now if you've got to this point and you still need to optimize, what we've done is just run through the different collectors under load. One of our JVMs is a message broker (ActiveMQ 5.16.x) and we push a couple thousand messages per second through it. We found that Shenandoah actually improved latency for our particular use case, which was more important that throughput and outright performance.
Oh, and if your application and usecase is _extremely_ sensitive to latency, forget everything I wrote and contact Azul systems about their Prime collector. They're pretty awesome folks.
And what is better strategy to set it?
I think basically the argument is that by setting the min bound lower, you allow the JVM to shrink the heap. This could maybe be beneficial towards reducing pause time because the JVM has less memory to manage overall. That being said, that SO answer also mentions:
> Sometimes this behavior doesn't give you the performance benefits you'd expect and in those cases it's best to set mx == ms.
I've also seen apps configured this way professionally for similar reasons. You might imagine some app that leads to the JVM pathologically trimming the heap in a way that isn't desirable and thus impacts performance in some subtle way, etc. The answer with a lot of this stuff is usually try both ways, measure, see if you can observe a meaningful difference for your apps/workload for typical/peak traffic.
Background: I've worked on Java apps for a few years at reasonable scale and worked on GC pressure issues in that time.
and now, rant time: well, Xms allegedly improves start times. Is that really important? No. Is that really true anyway? Not really. Xms hides problems... Yes. Give that memory to the OS. Let it fill it up with things like disk cache or COW artifacts. Xms is an attempt to help people that aren't planning properly or testing. Yes. Instead, Test your systems under full load with Xms off and adjust, measure, experiment, repeat.
Elastic is sort of a special case because it's a 'database', but, I'd rather know the minimum Xmx my system actually needs by experimentation, and you can't find that with Xms enabled. And even then, I don't see MySQL allocating 100% of its InnoDb bufferpools at startup.... :)
If you prefer high throughput, then relax the pause-time goal by using -XX:MaxGCPauseMillis or provide a larger heap. If latency is the main requirement, then modify the pause-time target. Avoid limiting the young generation size to particular values by using options like -Xmn, -XX:NewRatio and others because the young generation size is the main means for G1 to allow it to meet the pause-time. Setting the young generation size to a single value overrides and practically disables pause-time control. [1]
Setting -Xms and -Xmx to the same value increases predictability by removing the most important sizing decision from the virtual machine. However, the virtual machine is then unable to compensate if you make a poor choice. [1]
[1] https://docs.oracle.com/en/java/javase/19/gctuning/introduct...> Test your systems under full load with Xms off and adjust, measure, experiment, repeat.
precisely what I'm saying.
This was an eye-opener for me; That GCs performance do not depend on the size of garbage but on the size of the live objects, finalizers excluded.
And best part, every month there is a new bug released in regards to NMT reporting stupid memory usage that is hard to track coz regular metrics dont expose it.
So you have to enable NMT which in most cases is a straight 5-10% performance degradation.
And for latency use ZGC or Shanondoah.
And dont forget 50 jvm flags to tweak memory usage of various parts of memory that can negatively impact your production like caches, symbol tables, metaspace, thread caches, buffer caches and other.
God I hate that so much. Just let me set one param for memory and lets be over with it.
throughput: parallel or G1
balance between latency, footprint and throughput: G1
latency more important than throughput or footprint: ZGC or shenandoah
missiles and HFT: Epsilon
My java/kotlin app needs to keep a big table fully in memory. (~10 million records). And it is reloaded about 3 times a day. In C, I would just malloc the whole table in one chunk. Perhaps there is a specialized GC for this usage ?
For best performance you can decompose your structure to primitive fields (int, float, char, etc) and create array for every field. So you have, say, 10 arrays with million items each. Instead of creating one array which holds pointers to another 10 million objects on the heap. It gets tricky with strings (you need to flatten all strings into a giant char[] array and keep two arrays with index and length data, but doable.
Though 10 million of records might be OK for JVM. Measure your GC times.
I'd suggest to hide implementation details behind API, start with ArrayList<MyRecord> and refactor it later if needed.
Beware that you're gonna have to filter a lot! There's patch merging and very low detail implementation conversations. For example, you'd be pleased to know that G1 can now skip a guard in card-table clearing [4]. Don't ask me what is the card table and guards and why do you need to clear it, though.
One thing I'm looking for in GC advances is new hardware support for it in RISC-V J extension. There's gonna be memory tagging (helping security and memory management in GCs), and pointer masking (hardware support for what ZGC does under the hood)[5]. But we're probably a good 5 years away from seeing that in real life, if ever.
[1] https://mail.openjdk.org/pipermail/loom-dev/
[2] https://mail.openjdk.org/pipermail/hotspot-gc-dev/
[3] https://mail.openjdk.org/mailman/listinfo
[4] https://mail.openjdk.org/pipermail/hotspot-gc-dev/2023-March...
[5] https://github.com/riscv/riscv-j-extension/blob/master/point...
Besides the JVM resources linked by others, I found Richard Jones's "Garbage Collection Handbook" to be a decent introduction for background [1]. The Go team has written a bunch about their GC approach [2] - it is really interesting to see how it compares to the various options in the JVM and under which scenario you might prefer one or the other. And occasionally there are interesting articles on arxiv.
This leaves "full stack JS" in an awkward middle ground. Sure, you could still use it on your backend (like PHP), but why?
We use JS because there were literally no other practical options, but better browsers and WASM are providing new options.
Only you'll get higher-level abstractions (pattern matching, FP-style maps/flatMaps&Co throughout the standard library) and powerful collections with full support for generics.
Better yet, start with Kotlin from day one.
If you are looking for all the advanced type system features, kotlin is definitely not going to put you in the same spot, but if you are hiring Java devs, training them is a far easier lift: They can mostly train themselves.
I'd hope that after another version or two of scala3, when more organizations are happy running it in production, we might have an easier onboarding road, where we don't have to explain implicit parameters, implicit conversions and implicit classes, just so that we can get to the real meat that is the mechanics of type classes. But, as is, there's good defensible reasons for many teams to go try Kotlin first.
If I were starting a new company today I would certainly seriously consider Java or Kotlin instead.
Personally I would never trust JavaScript on backend.
:-/
TLAB is thread-local allocation buffer.
I think there's just a tpyo in the post.
During GC phases, threads can use PLABs to reduce need for synchronisation while scanning TLABs.
Java is such a storied and long-running and used-almost-everywhere language especially in Data Engineering (see all the Apache Data Eng projects like Calcite, Hudi, etc) but I just find it soooooo verbose and everything being a class and having to override things ugh .. it's all the things I hate about OOP in the forefront.
[1] https://openjdk.org/projects/amber/
Also, learning a language with new idioms is always worth it, regardless of whether you end up using it.
When it comes to super complex just look at any of the type signatures of the standard library for collections:
def ++[B >: A, That](that: GenTraversableOnce[B])(implicit bf: CanBuildFrom[IndexedSeq[A], B, That]): That
Comparing this to Java it is "super complex". IntelliJ can't even figure out the types sometimes.re: Complexity - At least the signatures for the core collections have been cleaned up a fair amount. That said, the richness of the type system and the prevalence of operator overloading always made it feel like a language you could be really productive in once you knew the language and the current codebase really well, but was really hard to just read through unfamiliar code and know what is going on.
How is this any worse than Java? My most vexing dependency-hell issues have involved breaking API changes to Hamcrest matchers and Apache Http Client; more recently Jackson-databind. All of those are Java libraries, brought in via transitive dependencies, usually from Java libraries.
> def ++[B >: A, That](that: GenTraversableOnce[B])(implicit bf: CanBuildFrom[IndexedSeq[A], B, That]): That
I should point out that this is the Scala 2.12 and and earlier signature. `CanBuildFrom` is gone from the Scala collections since Scala 2.13. In fact, the collections were redesigned for Scala 2.13 primarily to simplify method signatures, following community feedback.
Also, the reason why Scala’s collection has such complex signatures is because it is hands down the best collection lib out of any language I have used.
More features in Java (the language) gives other JVM languages a larger set of tools to design interop support around, while letting them remain a place for these features to incubate without the headache of the JEP. Sometimes these language-level changes may come with modifications to the JVM to support them as well, letting other languages clean up their implementations.
The whole situation works pretty well IMO.
Things like copying or creating derived records are a huge pain (or slight pain with code generators) while other languages have solved this long ago (even JS and C#).
This is the vague plan.
The reality is Java truly shows its age. Its dogmatic insistance on pure OOP and the enormous numbers of horrifyingly unintuitive design patterns adopted by the community cause even the most well intentioned engineering culture to eventually produce ugly codebases.
At this point if you're stuck in the JVM ecosystem you're almost certainly better off with Scala (which indeed is actually used) or the newer Kotlin.
In a way, for all the shit people give C++ modern C++ codebases are actually quite pleasant and the language is very flexible. Would I encourage its use for general production systems? Not really. That crowns is Golang's.
As a rebranded Limbo with Oberon-2 method syntax in 2009, not so much.
Classic essay being http://steve-yegge.blogspot.com/2006/03/execution-in-kingdom....
> nice and easy to understand
Java is definitely more practical than pretty. This frustrates the trendy crowd, but I think its important for tools to be practical. Shiny, pretty languages are never as long-lived as practical "ugly" languages.
I was one of the Bored Scala Crowd in the audience of a presentation by jetbrains people that must have been not too long after 1.0, took me half a decade to accept that a "poor man's scala" is actually the language that I want.
Regardless of programming language, every substantial software development needs to make serious architectural choices, otherwise things will become messy. I learned this the hard in my first programming job out of university. We were a C/C++ shop doing distributed software and one question to decide right upfront is: shared memory vs message passing. Both have advantages and disadvantages, but don't mix well. We started out with message passing, but over time somebody added POSIX shared memory as a performance hack (which two processes happen to run on the same machine). Over time, we added a second, non-POSIX shared memory layer for performance reasons under Windows. It was an unmaintainable mess. The core problem was that the CTO didn't understand distributed programming well enough to push back against those performance hacks. "How do you handle software architectural leadership?" used to be one of my replies to the inevitable "have you got questions for us?" in job interviews, when I was being interviewed. Now that I am making architectural decisions myself, I put a lot of efforts into this, for example building suitable linters that flag violations in code reviews.
If your team fights over cats vs scalaz or over Scala-as-OO vs Scala-as-FP then that's a sign that the technical / architectural leadership is weak. Choices like pure-OO vs pure-FP vs mixing them require care, and are hard to change in-flight. If you have strong modularisation, you can successfully use different paradigms in different modules, but strong modularisation needs care and architectural discipline, too. Part of the problem is that the style of programming that libraries like cats require (monads, functors, applicative, type-classes, HKTs) is not yet widely understood. This approach to programming needs its 'map/reduce moment' and become part of the mainstream introduction-to-programming curriculum. We'll get there, maybe in about a decade.
In my (considerable) experience of teaching programming, imperative programmers find jumping straight from imperative programmers to a monadic style extremely difficult, the more experienced, the more difficult.
Kotlin is worth picking up, but after seeing the speed at which Kotlin moves compared to how Java moves now, I don't think Kotlin will keep up long term.
Most of the time the complexity in my code has little to do with Java being verbose and more due to the business problem. There are areas to improve and Java's been making great strides recently. For example, in Java 21 we may finally have methods like getFirst() and getLast() for lists (via JEP 431) instead of the incredibly clunky list.get(list.size() - 1). Java also recently added multi-line Strings and templating is coming shortly. Streams and Optionals also reduce quite a bit of boilerplate, e.g. Optional's map and ifPresent methods are often elegant. Really I can't think of many other areas where Java gets in the way. Our team is incredibly productive with modern Java.
I think most developers actually write overly verbose code regardless of the language. And it seems little to do with years experience. This youtube channel covers most of the basics:
https://www.youtube.com/@CodeAesthetic
To me I just follow these recommendations naturally but in most PRs I review there's often huge amounts of overly nested code, poorly named methods, etc.
BTW, is that your channel?
I miss the data oriented approach Clojure libraries take. I remember having an issue with Ring and I just dove into the Ring source code and it was so simple and clear and I found the solution to my issue in minutes. I've never had that experience with Java and instead resort to forums, stackoverflow, etc. Most Java libraries would have several layers of abstraction and they're often intimidating. I consider myself an amateur Clojure dev and yet still contributed some PRs to some projects.
With that said, I've written some internal tools in Clojure and it was a nightmare whenever I had to modify them. They were pretty simple CLI tools that only needed updates every 6 months or so. I usually have to spin up the project with a repl just to understand the inputs and outputs to functions. I've ported all internal tools I created at my current company from Clojure to Java and I personally find them so much easier to maintain.