JRE 8 needs more codecache than before
engineering.indeedblog.com
engineering.indeedblog.com
That said 250MB is really a lot of code. Their binaries must be gigantic. This by itself likely already causes performance problems because all caches will be thrashing.
The issue has been three since 6.0.23 when the -XX:+TieredCompilation defaulted to true. (hmm, or was it 6.0.25?)
I'd say nothing really new. If you run a server, just disable tired compilation altogether. It might cost few seconds slower startup but that's it.
$ java -server -XX:+UnlockDiagnosticVMOptions -XX:+PrintFlagsFinal -version|grep TieredCompilation
bool TieredCompilation = false {pd product}
java version "1.7.0_55"
Java(TM) SE Runtime Environment (build 1.7.0_55-b13)
Java HotSpot(TM) 64-Bit Server VM (build 24.55-b03, mixed mode)
I would be interested in hearing from people with first-hand experience disabling TieredCompilation on Java 8 on production servers.After getting a baseline measurement of performance with JVM defaults, experiment with tuning a couple of settings, measure some more, make sure nothing breaks under load. Repeat until new JVM version seems stable/performs better/etc even under worst load and then upgrade all nodes to the latest and greatest?
Most of the settings found in legacy projects or on the internet are either obsolete or defaults in the latest version of the JVM.
The only mandatory heap flag is: "-Xms???M -Xmx???M" which set the heap size of the application to ??? megabytes.
It's only the server apps running 24/7 on fixed hardware that should have a defined memory usage.
The downside is, when we do need to use test data e.g. when there's a large change to the structure of the upstream or we want to do stress scenarios, it's a pain in the backside.
Some other techniques we use are one-box deployments that receive a proportion of production traffic and "bake" new changes before deploying to the whole fleet, and shadow fleets which let you tune and test against live traffic. We've found that simply replaying production traffic at higher volumes sometimes isn't sufficient, because our calls don't necessarily scale that way (some of them scale with upstream traffic, some of them scale by downstream fleet sizes due to client caching).
We don't have a complete understanding of why this app uses so much more codecache than other apps we've switched to Java 8. Now that we expose the codecache size in Datadog, I may try plotting code size vs codecache size for a variety of our apps.
I have no idea how many other people actually use it though.
For example, Presto is a SQL query engine that generates code for each query (a SQL query is effectively a program), so it can need a lot of codecache depending on the query rate and concurrency.
Isn't the JVM more conservative about inlining already inlined methods? I can't find a cite right now, but I swear I've seen this when inspecting the compiler output.
An unfortunate thing about using non-default JVM settings is that you need to scrutinize them every time you switch Java versions. If I ever go through this again, I will know to pay more attention to every JVM setting we override.
Beginning with Java 8, instead of having the VM to magically choose between using the client JIT (c1: think V8) or the server JIT (c2: think gcc -O2), the default configuration is to run in so called "tiered mode", first starts with the interpreter (as usual) then c1 (here keep the profile info) then c2. Because the code is compiled twice, you need a twice bigger code cache.
From my own experience, tiered compilation is nice when you run something interactive like an IDE (IntelliJ IDEA) and useless when you run a server app. That's said i've never had to have a 250MB code cache.
Yes, I agree if C2 recompiles something C1 already compiled, the C1-compiled machine code should be freed from the codecache.
Now, sometimes optimizations can be "wrong", in the sense that something new has happened that invalidates an assumption used during second-tier compilation. Here's a real-world example: I have an interface and only one loaded class that implements that interface. The second-tier compiler will use this information to basically do this:
function doStuff(MyInterface a) {
if (a.getClass() != MyClass.class)
deoptimize();
// Do everything from here on assuming
// that a is an instance of MyClass.
// This includes inlining simple getter
// methods to pure field accesses, etc.
}
Now, when a second class that implements `MyInterface` is loaded and makes it to that function, it will hit the `deoptimize` branch and go back to first-tier or even interpreted mode. Eventually the function will be recompiled by both tiers with the new assumption -- two implementing classes.So, in the case of deoptimization, it can be a win to keep the first-tier code to fall-back to.