It would be interesting to see what other languages were considered for the use case, and which cost/benefit considerations made Java come out on top.
It would be interesting to see what other languages were considered for the use case, and which cost/benefit considerations made Java come out on top.
Notable companies who started off with PHP or Python or what have you, reached a stage in their growth where they had to either migrate, or else write their own in-house compilers to fork those languages and make them scalable. No one's ever had the need to do this with Java (although different vendors do produce their own tweaked builds in order to compete in the support space).
People generally assume that a virtual machine is a disadvantage over AOT compilation. And for the "To Do List" apps that so many students and hobbyists on these forums are writing, they are correct. But JIT compilation is vastly superior for long-running server side business processes. Which is the niche that Java owns (facing some competition from .NET, which works in the same manner). It's not sexy or beginner-friendly, so the level of discussion that it gets here is disproportionately low.
How is that possible? Assuming all things are equal AOT should always be better. The advantage of a virtual machine is being able to distribute the same files to different processors and operating systems. The advantage of a JIT is that it makes your virtual machine faster.... oh how I wish Python would have an official JIT, sigh.
But back to my question, if a JIT compiler takes bytecode and converts it to native code at runtime, and an AOT compiler takes source code and converts it to native code before being distributed, then how could JIT be faster? Is it that many people use AOT compilers in a generalized way instead of taking advantage of hardware specific optimizations? Or is the Java JIT compiler just so good at optimizing java byte code that it doesn't matter if it's compiled ahead of time or not?
Edit: I got too many answers, but thank you all. I didn't consider the extra optimizations possible with runtime data. Please keep replying if you think of other reasons.
I wouldn't word it as vastly superior, because if you have a JIT you're probably doing some kind of tradeoff (most of the times, memory usage?), but the JIT can optimize based on the runtime profile, where an AOT't program cannot.
> JIT compiler can see that a function has never been passed a null object
It cannot. It can only see that function has never been passed a null object so far.
And the cost of such optimization might be that edge cases are way slower than anticipated. But what if those edge cases are actually the only thing that provide value for your business?
Imagine you're building a monitoring system that alerts very rarely, but reaction time is really critical for you. Do you really want to JIT compiler to optimize the alerting branch out because it hasn't seen any alerts yet?
My bet is that JIT compiler rarely does those optimizations by default, because it cannot know which branches in your code contain business value. I don't need my code to be twice as fast on average but twice as slow on critical branches.
Now, what would be great if profiler would make suggestions about how to modify your code for performance based on usage. But those should be approved and commented by humans, because humans know whether those are valuable optimizations or a fad.
This is truly a nit pick.
> Imagine you're building a monitoring system that alerts very rarely, but reaction time is really critical for you. Do you really want to JIT compiler to optimize the alerting branch out because it hasn't seen any alerts yet?
Yes I do, because once the constraints are invalidated the the runtime system will instruct the JIT to recompile the code with the now uncommon branch [0].
[0] https://shipilev.net/jvm/anatomy-quarks/29-uncommon-traps/
The link you shared talks about skipping compilation for certain branches to speed up the JIT compilation. Such optimization makes startup faster, but it doesn't make code execution faster. AOT compiled languages don't have slow startup problem in the first place.
It's like saying paper books are worse than e-ink readers because you cannot charge paper books.
If you're just using constants in your code to set logging level, AOT compilation can do exactly the same optimizations.
When the log level is not final, branch prediction Will ensure JIT code is faster.
>>... so far.
>This is truly a nit pick.
It's not even a nitpick, it's just wrong - your original statement is true: you said "has never been". The "so far" reply you received mistakenly implies you said "will never be" (given that HotSpot will deoptimise+reoptimise if its assumptions are invalidated)
That said, the remainder of their point I think is reasonable in the extremely specific (I think to the point of being unrealistic, given this seems like a hand-optimised assembly level of perf importance they have alluded to) scenario they mentioned - the JIT optimiser can't know that the user wants to optimise an extremely rare branch for performance to the point where the overwhelmingly common case should be degraded.
Another useful jumping off point to explore for HotSpot's performance techniques is [1]
1: https://wiki.openjdk.org/display/HotSpot/PerformanceTechniqu...
In order to be able to optimize presumably dead branches you need a primitive that:
1. Can detect entering "dead" branches
2. Is faster than branch.
I'm not saying JIT optimizations are not possible, JIT compiler can totally choose to inline function call based on frequency or loop length. It can make "better" time/space complexity decisions (though "better" is still going to be controversial).
But optimizing branches because they were not called in runtime seems like common a myth. Like previous commenter who confused startup optimization with code optimization.
2) If there are no nulls, it is faster not to do a branch.
https://shipilev.net/jvm/anatomy-quarks/25-implicit-null-che...
Once again, we're talking about how JIT can be faster than AOT, not how to reduce JIT overhead. Those are two different types of optimization. You will never be faster than AOT just by reducing JIT overhead. Both links that were provided talk about JIT overhead alone.
This one talks about internal JVM optimizations. Not about optimizations JIT is capable of doing for your code.
What exactly are we even comparing when talking about NullPointerException? To my knowledge, most AOT-compiled languages don't even have NPE (not in Java sense at least). It's apples to oranges comparison.
If that assumption gets violated a JIT can deopt and adjust later. In AOT you can only make the assumptions based on what the code can statically tell you.
JIT can indeed do a lot of tricks to reduce JIT overhead. But you can't present that as a benefit of JIT over AOT: AOT doesn't have JIT overhead at all.
JIT can definitely trade space for time. I bet that JIT will inline/memoize certain calls based on call frequency.
The only thing I'm arguing against is that low call frequency (like zero branch executions) somehow provides room for optimization compared to AOT. The only optimizations you can do in this case are optimizations for JIT overhead itself.
Anything beyond that is simply not possible: you can't eliminate a branch and detect the fact that you were supposed to enter eliminated branch at the same time: the simplest mechanism that allows you to do that is branching!
The trick of catching the SEGV can also be used by an AOT compiler (https://llvm.org/docs/FaultMaps.html), but the AOT compiler would need profiling data to do this optimization that is more expensive to do if the variable x happens to be null often. Even if you have profile data for your application, and that profile data will correspond to your application profile of this particular run (on average), you can not handle the case when x != null for five hours and then is set to null for five hours. If you use an AOT compiler you would have exponential blow-up of code generations for combinations of variables that can be null (if you would try to compile all combinations), and you would basically reinvent a JIT compiler --- badly.
Theoretically, a JIT can do anything an AOT can do, but the reverse is not true.
The article is talking about optimizations the JIT is capable of doing for your code. It is well written, and although it is talking about code throwing a NullPointerException, I think you can see that the same optimization can be done for my example at the top (that is applicable to c++ as well). So my comparison is not apples to oranges.
But your observation that most compiled languages do not have NullPointerExceptions is interesting. It could very much be because that it is too expensive with an AOT, have you thought about that?
At some level, the check for the alert has to be done on every call. If it isn't being done, the one time there should be an alert it will be missed. If it is being done, where is it being done if not in the optimised JITted code, and how is that code even more optimised than the JITted version would be? Why not just put that optimisation in the JITted code?
All sortsa stuff can happen.
edit: just read kaba0 answer below and in fact it's possible via tlb miss, really cool
The primary thing here is that the hotter the code path the more optimized your code will be using a JIT (albeit compiled with a compiler which is slower) which is impossible with AOT (since we have a static binary compiled with -O2 or -O3 and that's it) also Java can take away the virtual dispatch if it finds a single implementation of interface or single concrete class of an abstract class which is not possible with c++ (where-in we'll always go thru the vtable which almost always resolves to a cache miss). So c++ gives you the control to choose if you want to pay the cost and if you want to pay the cost you always pay for it, but in java the runtime can be smart about it.
Essentially it boils down to runtime vs compile time optimizations - runtime definitely has a richer set of profiles & patterns to make a decision and hence can be faster by quite a bit.
C++ LTO makes devirtualization practical to the point that your browser is benefitting from it right now: https://news.ycombinator.com/item?id=17504370
It can even do that if there are multiple implementations, by adding checks to detect when it is necessary to recompile the code with looser assumptions, or by keeping around multiple versions of the compiled code.
It needs the ability to recompile code as assumptions are violated anyways, as “there’s only one implementation of this interface” can change when new JARs are loaded.
One example of an optimization that cannot be performed AOT is anything that can be done in response to "noticing" that every time you get a List, you actually have an ArrayList. AOT can't make the assumption, but a JIT can start doing things like specializing/inlining code the implementation in question.
The question is, how much effort are you going to put into the profiling to determine that that's an optimization worth making? The larger the program, the more difficult such optimizations are to find -- the example given here was a trivial one.
You could do all of this by writing machine code, or constructing a programmable logic array for it. But would you actually get more trading done that way, or would be get more profit turning your programmers onto other tasks rather than one that can be handled by a machine?
in practice there is. For example, the JVM can inline virtual calls to code that was loaded at runtime (even if that’s third party code for which you don’t have the source code).
To do that by hand, you more or less would have to write something like the JVM yourself.
That's one of the biggest advantages of profile-guided optimization that works at the bytecode/IR level. It can recognize that generic libraries are only ever called with a single concrete type, and then specialize them to that type, and then inline the data representations of the concrete data types you actually pass, and then inline all the accessors that access that data.
To do that AOT you need to rewrite the library, which usually defeats the purpose of having libraries in the first place.
Or as one other possible instance of the design space you use the approach of cpp templates or any of Zig's comptime features -- when linking against a library you explicitly instantiate code corresponding to the things you know at compile time, apply optimizations, and throw away anything that's unused. Recognizing that generic libraries are only called with a single concrete type is child's play in those contexts.
It's pretty trivial to prove this wrong. Return a different concrete type based on a config value that is loaded at startup. AOT can't save you here. We can keep moving goalposts but it shouldn't stretch the imagination that a program's runtime characteristics can't be exhaustively predicted at compiletime.
This is sort of like claiming that branch prediction isn't useful because can't you just throw away what's unused? No, obviously not, so if the branch exists, runtime analysis can find out which one is more likely to be called over time (and generics are just a branch at the type level).
In any event, I was just saying that AOT can definitely cross library boundaries, not anything near as strong as what you seem to be arguing against (it seems you believe I'm asserting something like AOT being superior to JIT in all cases).
There are certainly going to be tradeoffs that runtime analysis can make that AOT cannot make, even if the example I wrote is possible to fix at compiletime.
Fun fact, the linux kernel actually has something similar with self-modifying code, that will remove a conditional at runtime.
1. AOT compilers have less information about the program than JITs (they only see the program as code; JITs see the code and the current execution profile)
2. AOT compilers can only make optimisations that they can prove don't change semantics, while JITs can be more aggressive, making optimisations that are speculative, deoptimising if wrong.
Presumably not the whole thing needs to be that fast, only the part that actually handles the individual events. Java has the advantage of being an industry standard, or at least being very common in the financial sector.
So you can work in a single language while only having to write some un-idiomatic Java in places where performance is critical.
Also, primitive-only java programming is probably no worse than similarly low-level C/C++/etc.
Firstly: memory safety. You might not be allocating but you're benefiting from optimized bounds checks and type safety. If you're trading at HFT speeds then you need your program to throw an exception if something goes wrong, and not simply make a trade with a random bit of the heap sent to the exchange in the 'quantity' field. And you need it to be easily diagnosed when that happens, without taking your whole service offline.
Secondly: these programs do actually allocate sometimes. They just don't do it (much) in the hot path in prod. For instance they'll happily write normal Java code that runs at startup to load data files, and they'll happily write normal code on cold paths that aren't hit all that often. They may well do allocations in test mode for logging, test verification etc. Some HFT shops do even allocate in the hot paths but they just size the heap so it doesn't collect during the trading day, and then they run a GC once markets close.
Thirdly: HotSpot is a very easy way to get PGO which is normally considered to be a 10%-25% performance win. Getting it from C++ toolchains is possible, but hard, and Java has the advantage that if market conditions suddenly change and you start going down different codepaths it'll re-compile on the fly, whereas for C++ you'd have to wait for the next rebuild to re-establish an optimal hot path.
Fourthly: not allocating isn't actually all that weird or non-idiomatic. Quite a lot of apps use this "allocate up front, but not in the hot path" approach. Check out Mindustry, it's a game written in Java that hits a solid 60fps at minimum and can easily hit hundreds of fps. It's just silky smooth even on a non-gaming laptop. How? Doesn't allocate in the rendering loop. Object pooling is hardly radical, most apps with real time latency requirements do that even in C++.
HFT folks in any language tend to break idioms to wring every cycle they can out of a machine. I'd bet you'd be faster in well tuned C++ code, but why not asm then?
I'd contend that non-allocating superfast code in any language ends up looking kind of odd. High Frequency Trade programs have a lot in common with embedded microcontroller apps. Except they're also doing business transactions.
You would also, I guess have to choose any library dependencies very carefully as well. Or implement the functionality yourself, negating any advantage that might bring in speeding up development.