Build OpenJDK for a Nice Speedup
august.nagro.us
august.nagro.us
So while building OpenJDK with particular optimizations for your particular hardware might be worthwhile in some cases, it's not for the faint of heart, and it should be done with care and extensive testing. If you intend to deploy such a custom-built OpenJDK in production, you should strongly consider getting the JCK [1] (the Java TCK) and testing your build for conformance (for all the VM configurations you'll use: GC choice, compiler choice etc.).
A safer and easier way to get "free" performance speedups is to use the most recent JDK (currently 13).
Is it a goal to remove undefined behavior completely? Because I seem to recall certain architectural decisions being completely reliant on undefined behavior, such as signal handlers as implicit null checks…
But moving more of the implementation of the JVM (or all of it!) to Java is the real solution to all this in my opinion.
Ah.
> But moving more of the implementation of the JVM (or all of it!) to Java is the real solution to all this in my opinion.
And what, ship a Graal binary to cut the bootstrapping/startup step?
The article benchmarks - Ofast which is poorly named. It's really -Obroken-by-design. It'll be "faster" but completely break applications.
It also suggests using omit frame pointer which destroys debugability.
-march and -mtune are the parts that the the article title and intro actually suggest. While possible, I see no evidence that this matters. As I understand it the arch that Java is compiled with is not the same as the one that gets used for JIT compiling.
Except for the jvm, apparently. And everything else I've tried it with.
> It also suggests using omit frame pointer which destroys debugability.
Which is completely useless except for jvm developers.
> As I understand it the arch that Java is compiled with is not the same as the one that gets used for JIT compiling.
The performance of the compiler itself matters, not just the performance of the generated code, because, since it's a JIT, compiler code continues to run.
IIRC frame pointers were necessary for a Linux flame graph tool I used on the JVM.
Well, yeah it's an article about performance optimisations, and disabling this debug mechanism increases performance (or maybe it does - I haven't measured it myself.)
If you need to debug, don't optimise for performance this aggressively. Seems a reasonable tradeoff?
But I believe compiler developers are trying to keep -O3/-Os working as advertised by their literal meaning, especially when it's combined with -march and -mtune, the cases of performance degradation should be fewer by now. For example, by using the knowledge of the subarchitecture, compilers can optimizing the code for Intel's MicroFusion rather than performing useless loop unrolling that actually degrades performance.
In all benchmarks on Phoronix since GCC 4.9, they showed -Os is almost always slower than -O2 and -O3.
https://www.phoronix.com/scan.php?page=article&item=gcc_49_o...
Linux kernel used to prefer -Os at everywhere, but now it has -O2 and -O3 as well. I think there is a measurable performance improvement in benchmarks in some cases.
https://github.com/torvalds/linux/blob/15f5db60a13748f44e5a1...
I'm currently writing a path tracer as a side project, and -Ofast is sickeningly faster than -O2. Like 4-5 times faster. -O3 duplicates all the loops into two versions: one with AVX instructions chunking 8 iterations at a time, and a second scalar version that does the final 1-7 iterations. -Ofast is a little bit faster than -O3 because it generates the approximate rsqrtps and rcpps instructions.
However this is a special case. Most code isn't heavy on crunching massive quantities of floats, and much of the code that does isn't written in a way that gives the compiler the freedom to autovectorize your loops. And a surprisingly large amount of code is still compiled with MSVC which won't vectorize at all.
In gcc, -O2 is generally fastest for general purpose code, both faster than -Os and -O3. As an example of why -O2 is faster than -Os, -O2 will optimize signed integer division by a constant power of two into a handful of bitwise instructions, which are larger than a single (but much, much slower) division instruction. (signed integer division can't be replaced with a single bitshift because negative integers work differently. unsigned integer division can be replaced by a single bitshift.) People think this is a universal optimization, but it's a size tradeoff so -Os specifically eschews it.
However, my point for a general purpose language system like OpenJDK that favors multithreading or multiprocessing or server workloads, each core/thread and process/thread share the I-Cache/L1 cache. Address lines are still 64 bytes. Code for these systems rarely is expected to run in isolation. I tend to want to be a good neighbor and reduce code size when I can.
[0] https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html
https://cpp.godbolt.org/z/-zeMWT
(stop reading now if that's enough)
The -Os compilation is some bookkeeping and an idiv instruction, the -O2 compilation is bookkeeping, two shifts, and an and instruction. I don't know what they because I'm drunk sorry not sorry but the one that does the idiv instruction is both slower and smaller and that's on purpose. (you can tell the idiv one is smaller by clicking the 11010 button to display offsets. the starting instructions of both functions have the same offset, but the final instruction of the -Os compilation is substantially smaller than the -O2 compilation.)
The description of -Os which you linked is telling. It says it's -O2 without certain optimizations, and then it also says:
> It also enables -finline-functions, causes the compiler to tune for code size rather than execution speed, and performs further optimizations designed to reduce code size.
Somewhere buried in that tuning for code size rather than execution speed and further optimizations designed to reduce code size is an optimization that will replace bitwise magic with division statements. -Os does what it says on the tin. It makes your code small. It makes your code slow. It does so on purpose.
People think that -Os is a superset of -O1 and a subset -O2. It's neither. It is neither a superset of -O1 nor -O2, nor is it a subset of -O1 or -O2. There are speed optimizations that -Os adds to -O1 and there are size optimizations that -Os adds to both -O1 and -O2.
The point, I think, of -Os, is for embedded. If you have a size n PROM, and your code compiles to size n+1 with -O2, and if you apply -Os and it compiles to size n, that's a feature. -Os, in my opinion, ought to be uncompromising towards that goal. For better or for worse.
I'll be sober in the morning and can engage with you better then. Sorry.
Reads like it came right from the Gentoo ricing guide.
But the general rule-of-thumb is that -Ofast should not be used, unless you know what the program is doing and how the optimization affects it.
A more meaningful comparison is -O2 vs. -O3 vs. -O3 -march=native -mtune=broadwell. Or run the OpenJDK test suite with -Ofast and see whether there are failed tests.
The public binary distributions have to limit themselves to what X86-64 looked like when it first came out in 2003, which means they can't take advantage of any new instructions that were introduced in the past 16 years.
That sounds like a bug to me which should be reported.