GraalVM
graalvm.org
graalvm.org
1. A modern JIT compiler written in Java that takes bytecode and transforms it into machine code. There is a plan that it might someday replace HotSpot [1]. However, we are probably a couple of years away from this.
2. A native image compiler [2] that uses ahead-of-time compilation technology to produce executable binaries of class files. This means startup times and memory usage similar to a language such as go.
3. An abstract syntax tree interpreter called Truffle. Which allows you to easily implement languages on top of GraalVM. With the performance of compiled languages but using an interpreter. You can read more about this here [3].
There are also other features such as a LLVM bitcode engine called Sulong. And various PolyGlot functionality to support integration of whatever language you want. [4]
[1] https://jaxenter.com/openjdk-project-metropolis-137318.html
[2] https://www.graalvm.org/docs/reference-manual/native-image/
[3] https://www.beyondjava.net/truffle-compiler-compiler
[4] https://github.com/oracle/graal/blob/master/sulong/README.md
https://en.wikipedia.org/wiki/Partial_evaluation#Futamura_pr...
2019 https://news.ycombinator.com/item?id=20987946
2019 https://news.ycombinator.com/item?id=19885922
I'm fascinated by GraalVM, but I'm hesitant to take even one step down a path that leads to Oracle using my actual software implementation platform to shake me down for money. I'd actually rather take the development effort and performance hit on using gRPC or similar to get my Java and Ruby modules talking to each other than to get locked into a potentially weaponized version of the Java platform.
That was really sad to read. Oracle is really hurting the java language and ecosystem.
If anything, their biggest commercial error may be open sourcing too much. GraalVM EE is quite expensive for what it adds over the open source versions.
Even Linux Foundation's Zephyr is under Apache license.
When you give back your improvements to the community, under any license, you free yourself from maintaining a separate fork.
I think it is somewhat analogous to LLVM: Graal exposes a "language agnostic" API which allows multiple front-ends. There's a front-end for Ruby, JS, C++, others. But this API is higher level than LLVM: it knows about objects, and allows tricks like querying properties across language boundaries.
But I can't imagine what this API looks like! For example, graaljs exists: how do JS semantics get expressed in a Java VM? What does a getProperty instruction look like? How does it know when to invalidate inline caches?
See especially Assumption, CallTarget, and DynamicObject.
GraalVM is basically a JVM with a new compiler, the Graal compiler. Truffle is the language implementation framework for GraalVM and designed to build AST interpreter. The Graal compiler can optimize Truffle ASTs through partial evaluation (see "One VM to rule them all" paper [1]), and produces machine code directly without using Java bytecode as IR.
> how do JS semantics get expressed in a Java VM?
Graal.js comes with a parser for JavaScript code that generates a Truffle ASTs for a given JavaScript program. JS semantics are defined within the AST nodes (see [2]).
> What does a getProperty instruction look like?
I think getProperty is implemented in `CachedGetPropertyNode` [3], but I am not sure as there are multiple other property-related nodes.
> How does it know when to invalidate inline caches?
Unfortunately, Graal.js lacks documentation for `CachedGetPropertyNode`. But I encourage you to have a look at SimpleLanguage [4], a JS-like toy language and the reference language implementation for Truffle with proper documentation. [5] explains how reading properties and invalidating inline caches works and it's all done using the Truffle DSL. The `@Specialization` annotation [6] might be a good starting point if you want to learn how it works. You may also want to check out the docs on the GraalVM website (e.g. [7]).
[1] https://doi.org/10.1145/2509578.2509581
[2] https://github.com/graalvm/graaljs/tree/a36b978aa328a8da3a7f...
[3] https://github.com/graalvm/graaljs/blob/a36b978aa328a8da3a7f...
[4] https://github.com/graalvm/simplelanguage
[5] https://github.com/graalvm/simplelanguage/blob/a25e385dd8626...
[6] https://www.graalvm.org/truffle/javadoc/com/oracle/truffle/a...
[7] https://www.graalvm.org/docs/Truffle-Framework/user/README
You express JS semantics by writing an interpreter for JS in Java. You can't write it in any way you want. Truffle provides a class library for the construction of interpreters, in which you define nodes in a tree. Typically this will be something like the abstract syntax tree of your language but it doesn't have to be. Each node is a class you define, with an execute method (handwaving away some details here).
At the start, you parse source or binary code into a tree of these node objects and then call execute() on one of the root nodes. For example each method or function in your language might be an independent tree, and the root node would be the node at the start of the function. Execute then calls the execute methods of the sub nodes and combines the results together, as per any basic interpreter.
The node objects have fields that contain information about the program, for example:
"return a + 5"
might turn into 4 nodes, a return node, a + node, a variable read node and a constant numeric node which has a field containing 5.
The Truffle class library contains code that measures how often a root node is invoked. After a while some roots will get hot because they get invoked a lot. This method has become a hot spot.
What happens then is the Graal compiler starts compiling the execute method of the root node. It compiles in a different mode to how Java methods are normally compiled:
1. Any time the code reads a field, the compiler pretends it's a constant if it's been annotated with @CompilationFinal. Even if the field is a mutable variable, the compiler acts as if it's not, and will read the value of the field as a constant. This can then trigger constant folding and further optimisations.
2. Any method call is inlined. This proceeds recursively until everything is inlined, stopping only at methods marked as @TruffleBoundarys. The compiler ends up with a single huge method representing the entire interpreter contents of the guest language method. Any calls past a @TruffleBoundary are in effect calls into the language runtime.
Once this is done Graal starts optimising. After inlining the method may be enormous, however, a huge amount of the code in the interpreter can be removed by these optimisations.
Dynamic languages have behaviour too complex to fully compile to native code. The amount of code required would end up being enormous and very slow, as it'd constantly need to look up basic things, like whether you redefined what + means. Therefore Truffle supports a variety of techniques to make them run faster.
One is an Assumption object. You can create these in your nodes and check them in your execute methods. It's a boolean flag. When compiling the JITC assumes the assumption is true, and deletes any code that would have been called if it were false. Your execute methods can call a special method on an Assumption to set it to false however. When that happens HotSpot will de-optimise all the compiled methods and force the back to your Java interpreter.
Another is a transferToInterpreter() method. It's special and says, any code that could execute past the point where I call this method should not be compiled. It means you can keep code that handles obscure cases out of the compiled method. Again, if the compiled method would end up executing a transfer, a de-opt happens.
Another is specialisation. Whilst your interpreter executes it's allowed to change the node objects in the trees. For example because your interpreter observes that at that point in the code you only ever add numbers together, not numbers and strings. The node can do a type check and if it fails, de-optimise and swap itself back to a slower but more generic node. Truffle has a thing it calls a "DSL", really it's a bunch of Java annotations, that automates this whole process for you so you can define a template node class with a whole bunch of different execute methods. It then generates all the actual node classes and the code to do the type checks and swapping behind the scenes.
There's lots more in Truffle to do with supporting debuggers, profilers, language interop etc, but that's the gist of it.
All these techniques added together give you a high level API for building HotSpot or V8 style advanced speculating JITCs, with little more than a specially written interpreter. It's not entirely automatic, but it's far easier than any other framework out there.
GraalVM makes a closed world assumption, so it can do things like dead code elimination (for library functions that don’t get called), and apply static analysis to inline virtual method calls, perform constant propagation, and so on.
This lets it greatly outperform the JVM, at the expense of some rarely used functionality.
Basically, it’s like switching from the JIT to a C compiler’s -O3.
https://jaxenter.com/graalvm-chris-thalinger-interview-16307...
Outperform in what metric? Startup time? Granted. Anything else? Not so much.
>Basically, it’s like switching from the JIT to a C compiler’s -O3.
Yeah and that would be pretty bad (it's not a good analogy to begin with). "-O3" doesn't have anything that a JIT couldn't have. The only advantage is again, startup time. JIT compilation has otherwise only advantages over static compilcation, especially in highly polymorphic code, such as java. Static compilation for polymorphic code is a joke in terms of performance... And every C++ programmer should know that.
I didn't look at the spec, but I would assume that AOT is only the bootup and then the JIT will take over anyway. This means, they'd do some basic precompilation for fat bootup and then use JIT again to optimize the code further based on runtime analysis. Everything else would be a ridicolous step backwards in time and make AOT completely useless, except for some niche scenarios.
(For an example, see some of the optimizations done by TruffleRuby for things like `myArray.sort.first` - which it apparently optimizes by terminating the sort as soon as the first element is sorted to the front of the array... and all that without any special hints in the standard library. Please correct me if I’m wrong... it’s been a few years since I’ve read in depth about TruffleRuby. And granted this example isn’t Java, but I imagine there are great parallels there.)
this seems to be a related thread on twitter:
Startup time matters a lot especially for Java applications. The reason Java never got to the desktop (including browser) I believe was the startup time.
Startup time matters a lot also for micro-services.
A 2nd great benefit of GraalVM I think is it makes it easy to integrate programs written in different languages, say Node.js and Java for instance.
What do you mean by this? There are a lot of desktop Java applications...
* Pycharm * Datagrip * Charles * JDiskReport * TexturePacker * Android Studio and associated tools * Apache Directory Studio * Zed attack proxy
That's just stuff that I've used recently and on an ongoing basis... I feel like I see quite a bit more of it.
I think it's mainly the fact that it has always been very difficult to create executable files.
Running java requires you to install the given version of JRE instead of just downloading an application and starting it.
You could not fit it on a floppy
JITs can do great optimizations in theory. But in practice they have only so much time to do optimizations.
And startup time, binary distribution size, memory footprint matter too.
That being said, the people at Spring don't recommend it for production use yet. Here's what they have to say about it:
"While GraalVM is now GA, GraalVM native image feature which allows ahead-of-time compilation of Java applications into executable images is only available as an early adopter plugin, so we don't consider it production ready yet."
https://github.com/spring-projects/spring-framework/wiki/Gra...
Wasn't this feature already supported in JDK 14 with the jpackage tool?
Jpackage bundles a small Java VM (with only the features you use) together with your compiled bytecode into a single executable. When it runs it starts up the VM and executes bytecode on that VM exactly the same as it would be if you were to run a jar on a preexisting JDK/JRE installation.
Graal compiles your full app ahead of time into a native code binary. There is no bytecode/translation happening at runtime. That is why GraalVM advertises faster startup and lower memory footprint.
So there are essentially 3 ways to run JVM (Java/Scala/Kotlin/...) code:
* compile into bytecode jar -> requires existing VM runtime
* compile into bytecode jar + bundle VM runtime -> no dependencies required, runs as the previous option
* compile into native binary -> no dependencies required, runs native, starts and runs faster
This is incorrect. Jpackage creates an installer which will unpack the VM image and any other resources.
I've tried it with a small CLI tool and it actually doesn't work that well (at least on Windows).
This is my experience as well. After installation, my installed app just didn't do anything.
Shameless plug for one real-world usage: the search feature of my personal blog is built as a GraalVM native binary, running as a serverless app on AWS Lambda [3].
Disclaimer: I work for Red Hat, who sponsor the development of Quarkus
[2] https://quarkus.io/guides/
[3] https://www.morling.dev/blog/how-i-built-a-serverless-search...
That RAM consumption sounds definitely over the top; if you still have the context, logging an issue would be very welcomed. That said, there's many libraries enabled by Quarkus (see quarkus.io/guides/), so every essential functionality should be covered by now. Still a question of course whether your specific library in a given space (like FreeMarker vs. Quarkus Qute) already is supported. In any case, thanks for checking back in regularly, things might look better next time already, as the framework evolves rapidly.
To compound the inanity - Graal is also the name of a new compiler used in the newer JVMs, in addition to the name of a completely different kind of 'JRE/SDK'.
GraalVM allows for polygot execution of a number of languages: Java, Javascript, Python, and anything compiled to LLVM bitcode.
It runs them all 'side by side' so there's no translation barrier when interacting between languages.
GraalVM also comes with a native compiler that allows you to have much faster startup time, though it comes at the cost of not getting more advanced runtime optimisations.
As far as 'performance' I don't think there is anything fundamentally different form the newer JVMs.
It grew out of the MaximeVM project at Sun labs.
Other well known meta-circular JVM,were Squawk for Sun SPOT and Jikes RVM.
Then on top of that, it provides a LLVM like infrastructure to build compilers and language runtimes, all in memory safe language like Java.
Additionally there is a long term roadmap to increasingly replace C++ parts of OpenJDK with GraalVM code.
It’s a bit of a shame, because it’s really nice tech, but it will likely continue to struggle for adoption as long as it’s owned by Oracle. I wish they would spin it off; the same product, with more or less the same commercial model, would probably work very well in the market, as long as it lived far away from the the toxic Oracle-licensing nightmare.
https://www.graalvm.org/sdk/javadoc/org/graalvm/polyglot/Con...
Here are a couple of demos:
https://github.com/graalvm/graalvm-demos
I said "currently", because you currently have to use the host Java, that is the Java on top of which all other languages are implemented. The GraalVM team is working on Project Espresso, a Java written in Truffle. That will make polyglot programming with Java much more consistent with the rest of the languages. Project Espresso isn't public yet, but here's the latest update from the team working on it:
https://github.com/oracle/graal/issues/1656#issuecomment-664...
What I found interesting with this approach is that it allows you to do server side rendering without needing a separate server and without needing any third party dependencies aside from GraalVM itself.
0. https://github.com/JackBister/javalin-preact-archetype
1. https://github.com/JackBister/javalin-preact-archetype/blob/...
2. https://github.com/JackBister/javalin-preact-archetype/blob/...
They also had virtual machines and five nines before it was cool.
What's the long term vision? In theory can this thing become "better" than HotSpot, JIT and G1/Z/Shenandoah?
https://www.graalvm.org/docs/tools/vscode
and supports the LSP: