This was over 10 years ago, so things may have improved, but it would certainly give me pause about ever implementing anything time-critical in Java.
This was over 10 years ago, so things may have improved, but it would certainly give me pause about ever implementing anything time-critical in Java.
This is using the GC that ships with JDK 11. Future improvements, such as Shenandoah and ZGC, may be able to further drop the interruptions to 5ms on average.
The radio repeater is designed for real-time, safety-critical operations dealing with voice communications. Java is up for the job; like others have mentioned, it's more about the technical team and whether management understands what it takes to write and maintain an excellent code base than it is about Java versus C++.
Tick to trade latency for a modern software HFT stack needs to be under 10 _microseconds_, which is over three orders of magnitude more stringent.
They already just said they need 10 microseconds. 10 milliseconds is far far too slow.
Another relevant aspect is what happens when one hits GC pauses again. Will there be any tweaks left? Will they conflict with the previous ones? Using such high-level levers to fix performance problems in a specific component is great when it works, but when it doesn't you're out of options.
The biggest issues are that an engineer with the knowhow is expensive, a team of them is prohibitively expensive, a lot of companies outsource this work to different countries and make them sign onerous and dubiously-enforceable contracts, pay little, demand too much, and have incredibly unrealistic deadlines for the type of work that needs to happen.
You know how it is, if it never fails, you're throwing money out. That's the mentality of a lot of management types I've dealt with.
That's certainly one way to look at over five years of engineering effort, extensive systems testing, countless hours of automated regression test development, independent product quality verification, rigorous code reviews, risk assessment processes, mandatory four 9s call quality requirements, integration analysis of JVMs with deterministic GC, and so on.
Judge not a galaxy from a single star.
Doesn't Java itself say it shouldn't be used for anything like this?
(MISRA-C and other odd dialects excepted)
Source: I worked on several extremely high performance Java systems at Sun and also worked optimizing various Java servers at subsequent jobs based on my Sun experience.
At one company (after Sun) I took on a backend service which was doing GC pauses every 3-5 minutes, the code was a dumpster fire mess. After cleaning things up we didn't hit a single full GC in two months of production use (they redeployed every two months so that's as long as the process lasted).
Productivity is the point. Even while paying attention to GC, it's much faster to write good code in Java than to worry about every malloc() in C. I love C, but if I need to churn out high performance code quickly, Java is the choice.
IMO (possibly biased by how much I hated C++ in the 90s) it is easier to maintain clean code discipline in a large team with Java than it is with C++.
Definitely depends on your team of course.
A few other reasons:
The rest of the platform for review and analysis doesn’t have such stringent performance requirements and it would be nice to reuse code.
There also may be some good libraries that are jvm only that your program relies on.
Packing up a fat jar can be a lot easier, less complex and more reliable than building a binary.
Your dev team has a lot Java expertise.
Getting Java to not allocate is not as hard as it seems if you set out to do it from the start. And even if not it’s not impossible because the tooling is so good. You just have to know the right tricks and it can be easier to incrementally learn those tricks for a dev team than learn C++ and it’s ecosystem.
Allocate either locally on the stack or allocate for the life of the process so it never gets GC'd.
Of course, spend time profiling to make sure you got it. Figure out your redeployment cadence and that sets the ceiling for how much data can be an exception to these rules.
(e.g. if you redeploy nightly then it's fine to grow the memory consumption as long as you don't hit full GC in less than 24hrs.)
In a nutshell that's it.
Programmers used to say that the GC and its fuc*#=+ STW was a pure mess, but under the hood, they didn't realize (accept ?) in fact the code WAS the mess.
Don't (always) blame the tool !
Aws is slow. Not talking about the network.
Consider that large chunks of Google, Twitter, Amazon, Alibaba etc run on Java.
The other major complication was that this was doing a lot of complicated protocol conversion, so it wasn't enough for the core message processing & routing to be solid and leak-free, all the input parsers and output converters had to be too.
10 years has seen improvements to GC, but not eliminated the problem.
Cass 4.0 finally enables ZGC, so we'll see how that goes.
If you have stateless, you could do two JVMs and monitor the GC state, and reroute requests to a low-heap JVM and let the other clean up.
Garbage collection is hard.
I'm surprised that there aren't expert-level concepts like weak references and ways to sub-partition the heap, with say a new operator that can target a partition, so major GC doesn't have to do the whole heap. Although that starts getting into rust concepts and analysis.