GC progress from JDK 8 to JDK 17
kstefanj.github.io
kstefanj.github.io
Once upon a time, this indie Java game called Minecraft became the most successful game of all time.
But from the few minutes of research I just did, Java cannot be deployed to many commercially important systems
- Nintendo Switch
- PlayStation
- iOS
It appears Java is only still viable for Windows and Android, and the 1% Linux desktop market.There used to be the GCJ project which would in theory let you run Java anywhere you had a C/C++ compiler, but Oracle's litigiousness killed that because the Java[TM] "platform" must run the official Java[TM] bytecode.
It appears C# via Monogame lets you deploy to all desktops (Win/Mac/Linux), mobiles (iOS/Android), and consoles (PS/Switch/Xbox). So ironically C# seems to now be the "write once, run anywhere" fulfillment of the original Java promise.
[EDIT: grammar.]
There are some engines, frameworks: https://jmonkeyengine.org/, https://litiengine.com/, https://libgdx.com/, https://www.lwjgl.org/.
But I have no real experience with any of those.
LibGDX appears to support iOS/Android and HTML5 desktop browsers, but not consoles (Switch/PS/Xbox).
iOS is actually not that hard, I got it to work using RoboVM, but you have to do some research because the documentation is outdated. HTML doesn’t work with Kotlin because it compiles directly from Java, and because it compiles directly it probably has some extra quirks.
At that point why even bother with Java or native images.
https://gluonhq.com/products/mobile/
They are also quite happy to sponsor possible console ports.
Don't confuse Java having the fastest GC with Java being the fastest GC'd language (especially not in all situations)
I've always wondered, it's likely that Java has the fastest GC because it needs to have the fastest GC, otherwise it would be a bottleneck. Other popular languages probably don't depend as much on the performance of their memory allocation primitives.
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
I haven't tried C#/MonoGame yet but all this discussion is considerably warming me up to it. The recent extremely successful indie game Hades was made with MonoGame.
Microsoft in typical MS fashion makes MonoGame higher friction if you aren't using a Windows box for development with MSBuild. On Linux it appears you need to use Wine to run DirectX effect compilation, though once compiled it works on OpenGL backends:
https://docs.monogame.net/articles/getting_started/1_setting...
It's impossible.
I just spent 30 minutes trying to find a single non-Microsoft mirror for the .NET Core dependency. If MonoGame is your Spice Melange then Microsoft's servers are the planet Arrakis, the only source in the known universe.
On Ubuntu you need to add a Microsoft server as a repo, you can't just "apt-get install dotnet-sdk-6.0".
So, yeah, MonoGame is not for me. I'll stick with Godot or SDL2.
[DllImport("pcre2-8", EntryPoint = "pcre2_compile_8", CharSet = CharSet.Ansi)]
extern static IntPtr PcreCompile(string pattern, long length, uint options,
out int errorcode, out long erroroffset, IntPtr ccontext);https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Only the PiDigits benchmark uses "extern's"; only because it seems like a lazy port. Everything else is native F# code. It beats Java in all benchmarks expect for the binary-tree one. To see a functional language match or beat benchmark level Java on many cases, at least for me, feels kinda nice. Many other real world benchmarks in house (tech choice evaluations) and third party I've seen JVM vs .NET Core also show .NET usually coming out on top recently.
The .NET GC isn't as good as the Java one. But I feel that's because the cost/benefit of improving the CLR's GC is less than Java's so the work is put elsewhere. The language (C#/even F#) generates less garbage in the first place with typical code. Any GC improvements there probably don't have the same bang for buck as in Java where allocations IMO are more frequent in day to day coding.
Doing so cut down the number of memory allocations to _3 allocations_ from something like _2 * N allocations_ (N being the number of elements in the map).
The improvement in performance was apparently staggering when it came to lessening the GC pressure. (Sadly I couldn't find it with a quick google).
https://github.com/microsoft/referencesource/blob/master/msc...
https://github.com/ConfettiFX/The-Forge
The Forge is graphics only; audio, input, etc., have to be handled by something else and Hades used a custom C++ engine apparently.
It would be great if the whole cycle were that fast! But alas, there simply isn't enough memory bandwidth to GC 128GB of memory in 0.1 millisecond :)
The amortized cost of GC is more efficient that malloc/free, but it's traditionally not good for latency sensitive systems due to long pause times causing jank/dropped frames. Now that GC algorithms like ZGC have advanced to give sub ms pause times, you no longer have that worry.
Technically we didn't need ZGC for this to begin with. IBM launched metronome a while back for real-time systems, but it was never as widely available as ZGC.
In a game you should keep allocations during gameplay to a minimum anyways, malloc() is not O(1) is has variable runtime based upon the current layout of free memory. Additionally, long running malloc/free based applications have unfixable memory leaks due to memory fragmentation.
In my opinion there's very few cases where you should be dynamically allocating memory, and not using a garbage collector.
Well I guess that's not technically true, free() could very well be a quick operation if you have a single 128GB object.
F# can be used with C#/MonoGame which seems to run well everywhere so that's one route to functional programming gamedev. Another route appears to be to use SDL2 with a functional language that supports ANSI C bytecode runtime fallback. E.g. OCaml has a bytecode compiler and an ocamlrun runtime that can be compiled with an ANSI C compiler. But I don't know what the GC latency guarantees are for the bytecode runtime. OCaml's native low latency GC benefits from Jane Street's contributions because Jane Street uses OCaml for high frequency trading. But bytecode OCaml running on Nintendo Switch isn't the same thing as native Linux OCaml.
The toy demos are quite elegant, although the community is probably lacking a good scene editor ie Godot or Unity.
Edit: godot-haskell exists, but it's still a little edgy.
I've been doing this with a game I'm building that has a custom 3D OSM map renderer.
Funny how you frame Java "only" being available for some of the most popular platforms, which is billions of devices.
Although, maybe it doesn't count as I used mods like Optifine, which are made to ... replace said un-optimized code, but I thought it was a good showing for Minecraft and the JVM anyway.
But it can overwhelm itself. There's a equipment called Elytra you can wear to slide down the air, and if you use Rocket when sliding, you gain a huge momentum boost.
If you keep boosting on a server, being fast enough to challenge the serverside world-loading, you can crash the server.
Another big defect is it's rendering is deeply tied to cpu time, and the game itself has limitation. My recent experience with a 200+ mods & shader setup is that with RTX3080TI (also better cpu) and GTX980TI it runs at same fps (20~30).
You guess right. There are some few fan made mods (optifine, sodium/etc) that often improve performance by an order of magnitude, from tens of frames a second to hundreds.
It's a shame for Java since the increase in render distance makes the game much more immersive.
In my eyes, there are no truly viable options out there, mostly due to a lack of approachable GUI game development software or toolkits.
For example, compare the one option that comes close, jMonkeyEngine (https://jmonkeyengine.org/) to the likes of Unreal (https://www.unrealengine.com/en-US/) and Unity (https://unity.com/), or even Godot (https://godotengine.org/).
Sure, many out there enjoy developing games in a code first approach, or even writing their own engines (e.g. Randy, whose videos are pretty interesting and comedic: https://www.youtube.com/c/RandytheSequel/videos or https://www.youtube.com/c/RandallThomas/videos), but i'd argue that the success of an engine largely depends on the popularity that it gains, which is largely influenced by how easily approachable it is.
Java game development doesn't have such a tool or set of tools, even the activity on jMonkeyEngine's GitHub (https://github.com/jMonkeyEngine) is really low, compared to that of Godot (https://github.com/GodotEngine), even if the technologies themselves could be used to similar degrees of success in many situations.
Come to think of it, it would be nice to actually benchmark something like Unity (C#), Godot (C#), Godot (GDScript) and jMonkeyEngine (Java) in similar real world applications, to see how they fare, performance, resource usage and development speed wise.
My intuition tells me that Java would be faster than GDScript, which would make talking about its (and also Java's, and thus also C#'s) performance a moot point for many of the indie games out there, since GDScript's slowness doesn't prevent many wonderful games from being developed in Godot, here's their latest showcase reel: https://www.youtube.com/watch?v=iAceTF0yE7I
It would be particularly interesting from the perspective of someone working in a shop which still has lots of latency-sensitive-ish workloads running on JDK 8 with CMS!
JVM is pretty well optimized and it is much closer to raw C performance than most other popular languages. You could also say that C is bad because its performance plateaued a long time ago.
Hell, now Java has much better SIMD support than C, even as a high level language.
C doesn't need it, you can just call CPU instructions like functions. SIMD is just another kind of CPU instruction, so C supports it. That works in C since you have a direct view of the memory layout. It doesn't work in higher level languages where memory is abstracted away from you, in those you need the higher level concepts you are talking about in order to take advantage of SIMD.
I think that's what people consider "better" about the Java approach (at least I do). That's of course not to say that you cannot do all this in C as well, but I think having these capabilities in that portable way available in Java makes SIMD useable for a huge audience for the first time which didn't consider that a realistic option before.
(TLDR, ZGC is a huge improvement.)
I've worked on a couple projects where switching from CMS to G1 was a pretty big latency win. Most of the were pretty strongly request-response-based. Pretty quickly G1 would converge on having most of the regions being young, and, by the time G1 wanted to do a mixed collection, most of them would have no live objects and would be summarily killed.
If I could have only 1 new GC thing from Microsoft, it would be the ability to totally disable GC during the lifetime of a process. I don't even want to be able to turn it back on.
I have a lot of scenarios where I could get away with the cruise missile approach to garbage collection. No reason to keep things tidy if the whole world is gonna get vaporized after whatever activity completes. Why waste cycles cleaning things up when you could be providing less jitter to your users or otherwise processing more things per unit time?
GC.TryStartNoGCRegion Method
Attempts to disallow garbage collection during the execution of a critical path.
https://docs.microsoft.com/en-us/dotnet/api/system.gc.trysta...
https://jet-start.sh/blog/2020/06/23/jdk-gc-benchmarks-remat...
Moreover, are the significantly lower GC pause times help improve snappiness of GUI apps, such as IntelliJ IDEA?
Any reason you're using sbt for Java projects? That's an odd combination. Anyway, I believe Metals is getting close to supporting this particular use case if you want to keep using VS Code.
However, now it is really bogging down. I'm sure it doesn't help that I've got several IDEs open along with some monstrosity of a "microservice" framework running half a dozen docker services and DBs on an older Macbook Pro, but it's true that app and OS devs tend to use all increases in CPU and ram, and so absolute performance and battery life never seems to get better over the years, regardless of how much the underlying hardware itself has improved.
Also, see https://github.com/JetBrains/JetBrainsRuntime/releases/tag/j... for a JDK17 runtime for Intellij products. I've been running this too and it works great (but does require some tweaks to the vmoptions file).
We run all our production applications on JDK17 too, which is why I push to run everything on 17.
It is not like it restricts what you can build with it.
We may not know their exact decision process but it is incredibly ignorant to assume they are just doing this out of laziness or incompetence.
* They've fixed long-known bugs in the JVM
* They've added sub-pixel anti-aliasing and other visual rendering enhancements.
Both of these are clear improvements on the JVM, and very likely to be accepted upstream, but instead of upstreaming them, they decided to create a long-running fork and package it as their own runtime. And like all long-running forks, they're struggling to keep up.
So we are left with two choices: abandon the hope of a functional IDE that only works well on a customized runtime, or abandon the litany of performance, capability, security, and stability improvements in the JVM standard and reference JVM implementation.
Have you had a problem running IntelliJ IDEA?
I am pretty sure that if I was distributing a very, very complex Java application for a very, very wide variety of audience (like people who aren't even developers because IDEA also caters to these people) I would distribute it with JVM packaged.
And then the user should not care which JVM exactly is being distributed. You only need to care about the end result.
Let me know the last time you cared which version of JVM is used by any of the SAS products you subscribe to.
I don't know why you would assume that I am asking for them to not package a JDK at all. That's a fucking absurd assumption, and nothing i have said has even suggested that. All I'm asking for is for them to upstream their improvements and package their IDEs with OpenJDK by default so we don't have to choose between Jetbrains improvements or OpenJDK improvements.
For starters, sub-pixel aliasing is a standard part of the JDK and isn't a Jetbrains addition.
What exactly is stopping you from using a standard JDK17 build with Intellij? How is the IDE not functional for you when doing so? How is it improved when using the JDK17 runtime they have on their github page?
Standard Jetbrains Runtime - slow as fuck, missing a whole host of performance improvements, security updates, and bug fixes since the fork occurred back in JDK8/9 Era.
OpenJDK - ugly as sin, gui becomes glitchy, lose out on customer support channels (first thing they tell you is to use the standard jetbrains runtime)
Newer versions of Jetbrains Runtime - have to reinstall manually every time you update your IDEs. You also lose out on customer support, as they are beta software.
This isn't really a hard problem to solve: find a way to upstream your improvements to the JVM, and then package the upstreamed version. All of these problems go away, and you're even relieved of the burden of maintain a long running fork.
- The current JBR is JDK11, not JDK8. - You don't need to reinstall anything to use an alternative JDK with the IDE. Just set the IDEA_JDK envar to the JDK of choice. - Ugly as sin is an opinion. I use the IDE with linux on a hidpi display. Looks perfectly acceptable to me...even running with JDK18. - I just reported an issue while running with the JBR17 a few days ago. Jetbrains was perfectly responsive to my issue.
I just tried running idea with adoptopenjdk-17 on macos, and it failed to start up, so it doesn't seem that simple.
[0]: https://confluence.jetbrains.com/display/JBR/JetBrains+Runti...
Trying to use a newer JDK on some applications, like Intellij, may require adding entries to the idea64.vmoptions file to relax the module restrictions that were tightened in JDK 15 and 16, if that app hasn't been updated for those changes. Entries like this: --add-opens=jdk.jdi/com.sun.tools.jdi=ALL-UNNAMED might be needed.
See https://youtrack.jetbrains.com/issue/IDEA-261033 for entries that might be needed.
[1] https://github.com/JetBrains/intellij-community/commit/edf1c...
This is an annoying behavior on at least Intellij Idea (I don't know about other Intellij products). If you ever increased its memory limit (it's an option on one of the top-level menus, and it will also occasionally suggest increasing it if it ever notices it's using too much memory), it copies the idea64.vmoptions file which contains not only the memory limit (-Xmx), but all the JVM parameters, from the binary directory to your configuration directory, and so the JVM parameters on it will be kept forever, instead of being changed when you update the IDE. The fix is easy: find that copied idea64.vmoptions file, write down the memory limit on it, remove that file, restart the IDE, and go to that top-level menu option to set the limit again. This will copy an updated set of VM options to your configuration directory.
For example first time loading project settings takes 1+ seconds while next less than 0.5. This value might get lower later if JVM decides to optimize more some parts of UI application.
The JVM is improving with never before seen speed, has state of the art GCs, a very good JIT compiler, upcoming green threads that will make blocking code automagically unblocking and once Valhalla hits with value types, there really will be very few areas where java would not be applicable.
And on top of that, there is also Graal, which is a novel way to run and optimize mixed-language code bases.
Seriously. It might not be one of the hip new languages with unique features but it’s a very mature language that is well battle tested, has plenty of libraries available, and the talent pool is very deep for Java hires. It’s also something that stands the test of time —- write code in 1999 on an old Java version it will still likely run on that version today.
"Hard Realtime Garbage Collection in Modern Object Oriented Programming Languages."
https://www.amazon.com/Realtime-Collection-Oriented-Programm...
The author is one of the founders of Aicas real time JVM, https://www.aicas.com/wp/products-services/jamaicavm/
"Distributed, Embedded and Real-time Java Systems"
https://link.springer.com/book/10.1007/978-1-4419-8158-5
PTC is the other company alongside Aicas, that still sells real time Java systems, https://www.ptc.com/en/products/developer-tools/perc
IBM J9 has the evolution of the Metronome GC, https://www.researchgate.net/publication/220829995_The_Metro...
And it has extensions for value types, via packed object data, https://www.ibm.com/docs/en/sdk-java-technology/7.1?topic=ob...
Someone else already referred Azul, they also have extensions for value types, called object layouts, https://www.slideshare.net/AzulSystems/jvm-language-summit-o...
Then there are special flavours like microEJ or the Android Java snowflake.
It is written in Java because the original devs probably did not know anything better.
If you work on Android games for example would you get any kind of control over mem layout, irrespective of the stack you will be using?
class RGB { float r float g float b ... }
and then make an `ArrayList<RGB>` the list will end up using over 2x the memory (128 bytes for object headers, 64 bytes for pointers in the list) compared to a language with value types. What makes this even worse is that since you are using a list of pointers, you can't use SIMD for any of your computations, and accessing elements will be slow since the values won't be in cache.
Also, there is now a Vector API that let’s you use SIMD operations with configurable lane width (and a safe fallback to for loops for processors without the necessary instructions)
In particular for Java there's regularly dependent pointer loads, which are dreadfully slow on modern CPUs and also waste a significant amount of L1/L2 cache.
Java can be ridiculously fast when only primitives are used so I wouldn’t put a quite performant game engine past it, but yeah, it would perhaps not be my first choice for a new AAA game engine.
That said, C# has value types and has for years now. Which is hugely important for arrays, and a major, major missing piece to the Java performance puzzle. C#'s FFI is also way better than the disaster that is JNI, which also plays a role here.
While Valhalla is in the works, but it is, I quote, at least 3 PhD’s worth of knowledge combined to figure it out in a backwards compatible way with all the interactions of generics, etc.
I develop games as a hobby, and my understanding is that this is literally true, but a little misleading. Games are as diverse as other areas of software, but they're skewed toward the more complex and demanding end. I don't think it's that atypical for a moderately complex game by an indie studio to get to the point where memory layout is a concern, which I don't think is as true for web apps, desktop apps, or other areas of development. That being said, I think worrying about memory layout is more than a lot of games need.
But most games don't actually need this. I'd be more concerned about whether there's a JVM language you like over other non JVM options for a game.
I also wonder how much of the progress to “Sub (200) millisecond” latency target is due us just having faster machines. I honestly have no model to tie this to actual performance of my code. I guess it translates but not really sure how.
Not bagging on Java —— I am just surprised how inefficient an industrial strength GC can be. I understand why manual memory management still holds its own now.
I'm not expressing a view as to whether the overheads of manual memory management are typically comparable or not. I really don't know.
Anyways, I've got plenty of anecdotal evidence where Java apps take order of magnitude more memory than their close counterparts written in languages that use manual memory management. Not browsers, but things like benchmarking tools, webservers or even duplicate file finders (shameless plug: https://github.com/pkolaczk/fclones#benchmarks - there is one Java app there, see its memory use :D)
The point I was making was just that you do need to compare empirically the typical overheads of each to make a meaningful comparison. The 6% figure in isolation doesn’t tell us very much.
As you point out, it is difficult to make these comparisons on the basis of anything other than anecdotal evidence, since it is rare for applications of significant size or complexity to be implemented in multiple languages.
?? These benchmarks keep the machine constant. And I’m not sure we’ve seen faster machines in a long time. Clock speeds have remained fairly constant and gains are solely in more cores.
Instructions per clock have risen steadily, even though the frequency stays the same. About doubled since 2011 according to cinebench[0].
That being said, I agree the faster machines argument doesn't hold much water.
[0]: https://cpugrade.com/a/i/articles/cbr15-ipc-comparison.png
https://cpugrade.com/articles/cinebench-r15-ipc-comparison-g...
6% memory overhead while preserving throughput is actually quite excellent. Just a little more than the average fragmentation overhead under manual memory management.
Manual memory management can "hold its own" because it can tailor the allocation/release profile to the problem and aggregate some of those overheads.
Also, Java’s GCs are compacting, putting similarly old objects relatively close to each other. The same random chat program in C might use a linked list with worse characteristics so it is really not that obvious to me what would be a good solution.
I suspect the answer is that ZGC is not considered mature enough, and has a higher memory overhead than G1.
It would be interesting to see how much of the improvement is just down to the use of extra "unused" cores, and how much CPU is actually used by the GC. Equivalently, run a CPU-bound task on one core and measure GC and application performance, with an eye on how much the fancy GC slows down the program.
That’s… not incompatible with being useable in a low latency service, but 100us is 2.5 x longer than a sync disk write these days, and they don’t report max latency.
My take on the article is that GC has gotten quantifiably better since Java 8, but not qualitatively better:
The types of projects that had to abandon Java due to GC (and JIT) latency 10 (or 20) years ago still shouldn’t consider using Java.
Like, yeah, I would probably not write something like pipewire (linux’s new audio processing project) in Java, but other than that and like some really low latency trading niche (where FPGAs are dominant now as even general CPUs are slow), where would it preclude the usage of Java?
So even if there's a reason to change, people wont do it.