> Oracle plans to contribute the most applicable portions of the GraalVM just-in-time (JIT) compiler and Native Image. Oracle does not currently intend to contribute the polyglot technologies supporting other languages such as Python, Ruby, R, and JavaScript. Additional details will follow in the coming months as we move forward through this process.
It would appear this is about making native image artifacts / distribution a first class citizen across all of Java - making it an alternative to uberjars + jvm for running/distribution. Ie native desktop apps and native binaries for servers?
Does that mean to run it you need another JVM to run it on top of? That sounds stupid... Maybe you need another VM just to run it on once, so it can translate itself to native code on the target?
GraalVM is the regular HotspotVM integrated with the GraalVM compiler (which is normally AOT compiled as a native library, but can be run as a JAR); it also supports tooling with the ability to compile code targeting the JVM to a native executable, as well as leveraging AOT compilation itself.
Graal reuses all the insanely good GC implementations, observability tools etc already present in OpenJDK (as throwing all that away would be stupid).
I haven't looked at it for ages and looking at Wikipedia it seems dead, so curious if there is any newer attempts....
You're commenting on an article about a newer attempt?
As pointed out by other comments, this concept of a "meta-circular VM" isn't new. It's been done before by two other projects, Maxine and Jikes. What's different about SubstrateVM is that this is a production tool rather than a research project, and it's not just a JVM, it's also a way to pre-initialize the app. Therefore programs compiled with native-image can start as fast as programs written in C. Actually, slightly faster in some cases. You may wonder how that's possible given that Java apps normally start slowly, but it's because there's no JIT compilation and the state of the heap is snapshotted, with classes pre-initialized including the JVM itself. So the program can literally just start executing at main() in machine code with no VM startup overhead, because it's done already.
The downside is that snapshotting and AOT consume a lot of disk space.
There are some questions below asking how this works. It sounds initially "impossible", like a lot of stuff GraalVM/Truffle does, but it's quite easy to understand really.
You start with a bytecode compiler written in Java. This is a normal program written in the normal way, because a compiler is ultimately just a function that converts one stream of bytes to another. Then you write the runtime and GC in Java too, and compile that as well. This code is a bit special. It's still syntactically Java, but, some classes and methods are given special meanings and some extra rules apply. They aren't compiled in the same way as normal Java code. For example you can write code like this:
UnsignedWord value = Pointer.readUnsignedWord(address)
This doesn't allocate an object or call a static method. Instead it will be compiled down to a single mov instruction. Likewise for writing to memory - there are magic methods that are taken to mean "emit this assembly" instead of doing normal method calls.Several other tricks are required. GC code can't allocate because it would mess up the heap it's working with, so you can use annotations to mark methods as "never access the heap". But then, GC code is written in Java and Java must allocate for almost anything non trivial, so how does that work? The answer is, the GC code is initialized at build time and all the objects it needs are snapshotted into the default heap that's mapped into memory at startup.
There are lots of other tricks, mostly annotations that control the compiler so that e.g. methods are guaranteed to be inlined and removed, objects are guaranteed to be stack allocated. This isn't available to normal Java but when you control the compiler it's not a problem. The advantage of this Java-superset (or subset) is that you can use all the normal tools that understand source code, like IntelliJ, JavaDoc etc.
Also, may I ask how do you know so much about the topic? I would really like to one day work on OpenJDK/Graal, but I just don’t see the road ahead me.. — I’ve just started my master in CS, but I don’t feel it closing the gap at all. Surely I can read up more and more on the topic in small steps, but I would be very grateful for any guidance/pointer.
How did I learn about it - mostly by reading their papers, watching their videos and asking lots of inane questions on their Slack. Also, I happen to live around the corner from where the Graal team work so occasionally I've been able to meet them in person and ask questions then. But mostly I just followed their efforts for a long time. I got interested in Graal back before most people had heard about it, after somehow randomly encountering a discussion of TruffleRuby on Chris Seaton's blog. Then I wrote about it here:
https://blog.plan99.net/graal-truffle-134d8f28fb69
Most of the Graal guys came out of masters and PhD programs at JKU Linz, so the path you're on is a well trodden one. I wouldn't feel down about it. For me, how it worked was quite mysterious for a long time and then one day it clicked, and I saw the essential simplicity behind the concept.
It's not uncommon to do a CS master Thesis as part of an internship, if your university allows that. We have plenty of topics to choose from and you can also come up with your own.
Good luck!
And (shameless self plug), implement a Truffle language. Start here: https://www.youtube.com/watch?v=pksRrON5XfU&list=PLN193mR3Js...
For example the Go runtime and I think C# as well are written in their respective language.
I can't quite imagine how you'd bootstrap something like that.
See Scheme48 or Go. There are some ways around that bootstrap problem.
That's what it does.
With Graal, it's possible to write a JVM in Java, but the JVM doesn't depend on another JVM to run, and the way it runs bytecode isn't the way it was compiled in the first place. It's not really self-hosted in the same way that a compiler can be.
It's like how PyPy is Python but with the asterisk that it's bootstrapped with RPython which is an almost-subset of python so that it doesn't require a runtime.
You could define a statically compilable Java subset (like that which gcj used to accept) and build a runtime in that which would mean omitting features such as reflection but a lot of defacto standard java tooling like Spring Framework would not be compatible.
But the JVM you build using your statically compilable Java subset can then run the Spring Framework or whatever.
How restrictive do you think the subset is? It only doesn't support some features you probably never wanted to use anyway, and arbitrary reflection. I maintain 125k lines of Java that conforms to the subset rules, and to be honest I never even think twice about the fact that it's a subset.
There are a couple of Java implementations written in Java, Jikes RVM being one of the first ones, almost 15 years ago.
BTW, how do you implement a garbage collector with a garbage-collected language?
You write it carefully so the garbage collector itself doesn't also need to allocate objects.
Here's a GC for Java written in Java https://github.com/oracle/graal/tree/master/substratevm/src/....
Java doesn't provide a primitive to deallocate memory. So while I can see how for instance allocation a huge chunk / big array could be allocated and you represent objects in there don't you end up with a situation where your process will always occupy a fixed amount memory? Might not need to be fixed. You might also be able to extend more but how would you free that again?
But you might find this interesting as a specific example - this is where it actually obtains memory from the OS.
https://github.com/oracle/graal/blob/44e68777b130c8ee781c72b...
Note the @Uninterruptible annotation - that's saying that this code is safe to use within the GC itself. Notice how the file doesn't contain even a single 'new! (Outside of PosixVirtualMemoryProviderFeature, which is something else.)
You are writing a GC in Java, yes but have access to low-level memory abstractions/interfaces, right?
Whereas I was initially wondering how to write a GC in "pure Java" that doesn't have access to low-level memory interfaces.
Does it make sense why I was asking, now? Or am I still not getting it?
To be clear, Java the language is Java, I'm not going to argue that it's not Java because of special primitives / interfaces available to write the GC here, but it is not what most people would think of when considering the limitations of the runtime everyone's using.
I think I'd phrase it like this: you can write a GC mostly in Java.
Other GCs than the default Java one in native image like G1 actually embed the C++ version of the implementation instead of writing it in Java.
Yeah, but also the code shown above used native low level primitives like mmap which typically aren't available.
AOT, special behavior directives (@Uninterruptible), native memory access, making sure not to use new (?), at that point you are formally using Java, the language, but it's sort of its own thing.
So, that's still cool and likely no way around it but I wonder to what degree it's actually beneficial: Your program is much closer to a C++ program than a Java one except for syntax and the additional glue abstractions that typically exist in neither. In a (exaggerated) sense it's like it's written in a C++ DSL embedded in Java.
If you know how to write C++ it's possibly simpler to just write it in C++ as you know how memory management there works. If you know Java you need to get familiar to the extensions used here.
This is a bit different from the idea of self-hosting a compiler for instance where you can start writing your compiler now ideomatically in your own language instead of something different. For instance if your language & runtime has GC it's much simpler / safer to write a program in it, and so a compiler in it than pre-self-hosting (if the earlier compiler was written in C).
So what in particular is afforded by writing the GC in "Special Java" ? Is it just about being able to say "it's all in Java" or are there language features in "Special Java" that make live easier than C++? Or other benefits? For instance, I imagine use of the wider ecosystem / libraries isn't possible whereas in C++ it is (more so at least).
It would be nice if somehow you could write a GC itself in regular old Java without having to worry about aspects of how to do it in a special, restricted way. Say, the code you write and which runs the garbage collection creates objects, which subsequently get cleaned up by your very own GC program.
In the end you still have to work with low level abstractions / memory though and maybe that's still different to the self-hosted compiler example in a sense: There you translate code into a language that is more low level but you don't need to have these abstractions available in your own language / on your stage of computation / while running your compiler.
Native calls are a normal part of Java. And the pointer and word classes used could be implemented using the Unsafe class which is present in Java.
the gist is you treat memory as a big array, then write a program to manipulate that array. it's really just a decision about what you want to "take as primitive" in your implementation. could be brk, could be malloc and free, or something higher level.
My turn: NASDAQ moved from cpp to Java quite a few years ago. Do you think you know something they don't?
SweetHome 3D IS slower than competitive projects written in C++.
C++ tools tend to be faster than comparative Java tools because of fast startup time, no GC, and being the default choice for performance sensitive projects for decades.
C++ is definitively faster because of inherent advantages.
You appear to be jumping to the third based on the first which appears erroneous in argument even if you turned out to be correct. In actuality it appears that for most things in the same ballpark language choice isn't necessarily the only or even the most important factor. This is even more true for things where startup time is an inconsequential factor, with better GC that doesn't result in lengthy pauses, and where development time is a substantial limiting factor wherein being quicker to work with may result in more time available to improve other design choices yielding as good or better results.
By the way, the .NET CLR is AFAIK written in C++, like with Hotspot. It's not fully self-hosting.
It's a tricky process because HotSpot is highly performance sensitive code. People won't accept regressions just to convenience the JDK maintainers. Java meanwhile is deliberately a simple language to make it accessible for people, so you lose some low level techniques that are useful for performance. Nonetheless, GraalVM native image has proven that you can achieve HotSpot like performance with a JVM written in Java. However it requires a big change in the compilation model that isn't always appropriate.
Some of the reason is historical, and some has to do with warmup. Project Leyden and Graal's Native Image will help compile more Java AOT, allowing even more of the runtime to gradually be written in Java. It will take some time as it's not a top priority: it won't immediately deliver user-facing functionality, and most of the work on the JDK is already done in Java code anyway.
GraalVM is the evolution of MaximeVM, originally developed at SunLabs, also about 15 years ago.