Java on Truffle – Going Fully Metacircular
medium.com
medium.com
If you can read Java code and you are interested in how it is implemented start reading here: https://github.com/oracle/graal/blob/espresso/espresso/src/c...
If you are interested in a deep dive into the Truffle technology and are interested in why it makes sense to implement a language on top ofTruffle checkout my recent Rebase talk for a really long answer: https://youtu.be/O4icaN9khp4
If you don't have enough time for an 1h talk checkout my summary tweet: https://twitter.com/grashalm_/status/1212763944043139072
I think if there's ever room to move towards Apache 2 or similar license it would be really welcome.
Oracle v. Google has nothing to do with it (Java’s licence at the time explicitly disallowed mobile usage) Do you have an examples that give rise to any sort of concern about the usage of GPL-licenced code (regardless of code owner)?
That is alarmingly false. Do you have a source that McNealy indeed has said that? I couldn't find anything by googling.
Source: Official Reporters for the US District Court for the Northern District of California (regarding Oracle v. Google, 3:10-cv-03561)
That Tetris game is that a command line thing? If so what library are you using to manipulate the console? I've wanted to create interactive command line utilities with the JVM for a long time and it's always been not supported.
That line turns off buffering so that keypresses can immediately be read, by putting it into "raw" mode.
ncurses I think. There's an issue with how jconsole and tetris work together, so when tetris exits, jshell does too.
Isn't this only true for GCC code that operates above the GENERIC form level?
The ‘GC’ refers to garbage collector, not GCC.
It is true that you can write a compiler frontend targeting a generic backend (like GCC), and take advantage of existing optimizations. Truffle's somewhat unique value prop is that you can write a basic interpreter (not compiler) and get performance approaching that of a top-end JIT.
Oracle has a very large, well-funded, and long-term PL research lab that consistently produces top-tier papers, and funds and collaborates with many academics.
If it doesn't come to mind when thinking about PL research then I can't think you're reading much PL research.
Most DBA I met would not argue against Oracle having the best Database. Whether they would choose to use it for their projects is entirely different though. Something that is objectively good has zero to do with whether you like or dislike the company.
The amount of work that is going into JVM and VM research is absolutely insane. One VM to Rule them All, [1], that was submitted to HN in 2013.
And for what it’s worth, TruffleRuby on Truffle/GraalVM (reportedly derived from JRuby) is supposed to be incredibly fast.
Microsoft abandoned IronRuby and IronPython around 10 years ago. Meanwhile TruffleRuby on GraalVM is the fastest Ruby implementation by far.
From Microsoft itself, F**, F#, VB.NET, C#, Sing#, System C#, Managed C++, C++/CLI, Cω, Axum, and who knows what else they have done internally.
From Microsoft partners, COBOL, Eiffel, Ada, Fortran.
From FOSS community and university researchers, not so much because it is M$ and better not use Mono because of patents, so was the saying.
That means the DLR is effectively dead, I think -- along with all languages for which the CLR's static call sites mean slowed performance.
Plus on UNIX, Java has 25 years of history, not 4 (not counting WIP mono here).
I spent most of last year maintaining a Web Forms app. It's no fun. Ironically, there used to be many more CLR languages, because Microsoft paid for their development. IronPython and IronRuby stand out. I wish ClojureCLR weren't dead too.
A google search would tell you otherwise.
https://groups.google.com/g/clojure-clr/c/2v9xls5hreE
Last commit 6 days ago: Last code change 11 days ago: https://github.com/clojure/clojure-clr/commits/master
Is there any reason for saying that? Checking the commits history on Github[1], the project seems still alive.
Either way it's pretty clearly not a core platform for Clojure.
And if you tried running your own EAR files, it could take an hour from a code change to actually test it. Building it, installing it, reloading the server, and then getting to the correct state of testing. If you discovered a bug, you had to patch the code, build the file again and start over, wasting another hour. Some of it was evaded by using JRebel, but if the changes couldn't be swapped you were SOL.
No way WAS+EAR can be compared to running stand-alone images. It was a clusterfuck of dependencies and dependent configuration.
Yuck.
It mostly goes to show that containerizing application code for execution in a managed environment in nothing new. I'm sure this was also done somehow in the mainframe world.
As to whether K8S is a superior embodiment of the concept, I'd say that the problem is rather that a large majority of adopters are just cargo-culting without having a real need for what it actually provides, externalizing the costs of complexity it brings. I have nothing against K8S itself, but its adoption curve is telling of CV-driven architecture at it's worst. Mind you, the same might have been said of J2EE back in the day.
LPAR -- https://en.wikipedia.org/wiki/Logical_partition
Upon understanding this, I realised that there is / was nothing new under the sun.
https://en.wikipedia.org/wiki/Burroughs_large_systems
https://en.wikipedia.org/wiki/NEWP
If you follow NEWP manual from Unisys (nowadays still selling Burroughs as ClearPath), you will see UNSAFE code blocks and how its use tainted binaries, which requires admin permission for execution.
Yes, it's called an operating system :-)
The problem was when it was used for multiple services (as his comparison with EAR and docker implies). Lets say you had a service for something packaged as EAR. You would then have to log in to the dashboard, add the package and configure it manually. If something changed with the package, you would probably have to reinstall it in WAS, you couldn't just pull the newest from git in your repo and be done. And if that needed any config changes, everyone on the team had to do the same locally and manually in the control panel. When you got multiple moving parts, you would spend more time bringing them in and configuring everything correctly than actually developing whatever you were supposed to.
But what I wasted most time on was certificates. Everything had to be signed and stuff, mostly self signed by randoms and more theater than security. Sooo many times stopped working because the certificate for some module expired, or the CAS was unknown and had to be manually loaded into everyone's local install etc.
Just last December I was having similar fun with certificates and "modern" approaches.
> ... it already passes >99.99% of the Java Compatibility Kit ...
> Running HelloWorld on three nested layers of Espresso takes ~15 minutes.
Link: https://github.com/oracle/graal/tree/6dd83cd94763b8736b42063...
It’s like how time works in “Inception”.
> other mainstream languages
And indeed, there are a couple of things that Java and .NET still don't do.
If anything, XML with schema works out of the box in most IDEs with autocomplete, validation etc.,
The problem with XML is mostly syntactical: it's not always clear if something should be a node or an attribute, and it doesn't have actual lists, maps, or scalar types.
JSON has more structural issues: it can be too terse, so if you're not careful with naming fields, it can be really hard to tell what type of object something is, and for polymorphic lists, you're out of luck.
This past year I wrote a rather simple desktop app (JavaFX) for work using Clojure. But Java by default uses RTTI - and Clojure by extension does as well - the dependency tree would balloon my final executable. For instance, I'd bring in some CV library like BoofCV to do some very basic image manipulation, but all of BoofCV and its dependencies would get dragged in. Hundreds of megabytes of junk.. So you either have to just accept you have a gargantuan executable or you have to use some smaller subpar libraries to do what you want. But it seems unfortunate the ecosystem pushes you to use less/smaller dependencies instead of using robust large ones.
I understand the reason behind it. The java compiler has no way to ensure which classes will be used and which won't be b/c crazy things can happen at run time. Maybe with all the work on GraalVM people will move their libraries to not use RTTI as much? I haven't had a chance to hook up Graal yet (I think it doesn't play nice with JavaFX.. or maybe that's old news) - but it's unclear if it'll slim down the final binary. If it does - then this potentially a game changer for "deliverable" executables :)
PS: You can manually go in and disable sub-dependencies.. but it's ugly and incomplete. I got Clojure/JFX/Java to compile down to something like 60MB - but my app is simple and it really should be more like 2MB.
Yes it does whole-world analysis and only compiles in classes, methods, and fields, which are actually used.
The executables do have a minimum size, because the GC code for example doesn't get any smaller, but I think 2 MB is around the area of what can be achieved, yes.
I understand javascript does tree-shaking, mainly in the context of code-minifiers, and certain languages do "path optimisation" usually at runtime - but I'm not sure how to describe the removal, and modification[0] of code (say, interpreted code) outside JS.
I'd like to "tree-shake" python in order to see what % of code in it's deps is actually used.
[0] e.g removing static parameters, replacing fns with lots of cases into multiple fns etc.
Yes, everyone else calls this dead code elimination. [1]
(the custom term isn't completely unreasonable, since tree shaking is somewhat different from how compilers do DCE)
Would Oracle control the implementation of a particular language that used Truffle? Could they decide to "deplatform" that language? ("we don't authorize the use of Truffle in language implementation X for Y reasons") Or charge for licenses for users of that language? Or other somber scenarios... according to current licensing or in a hypothetical future ...
Compare with, say, LLVM infrastructure, were it seems more certain that a language can be implemented as free software and with less restrictions or caveats.
This happened with solaris: oracle decided to make new ‘official’ development closed, but the last opensource version continued to be used and developed by the community.
0. https://github.com/oracle/graal/blob/master/truffle/LICENSE....
Graal licensing seems pretty complicated to me, and if I get it right Truffle only makes sense when used together with Graal.
From what I could find in the FAQ [1]:
* "GraalVM Community is free to use for any purpose and comes with no strings attached, but also no guarantees or support."
* "GraalVM Enterprise includes Improved performance and security over GraalVM Community"
In general it seems pretty good, but I think it is a bit painful to have to pay for certain optimizations... again, imagine if projects had to pay to, say, "clang -O3". Also could be painful to get stuck in Graal/Truffle version N if version N+1 switches to different licensing terms.
In general commercial compilers are a bit of a drag to work with... I use Saxon from time to time and sometimes I need to check the matrix to see if a feature is available or not [2]. For instance, until version 10.0, higher order functions (!) was a feature one had to pay for, for whatever reason.
--
> they are really bad at shepharding, just look at the state of android java vs openjdk
This has nothing to do with copyright law―the thing that Google was sued for. There is no legal argument in this remark (which is the problem with about half the comments that appear saying that Google was in the wrong), just an assertion based on an appeal to emotion that Google deserved to be sued, and then working backwards from there to present a half-formed argument.
It had a specific license explicitly disallowing mobile use. Everything else is irrelevant - google knowingly broke the license, didn’t they? This is copyright infringement. As for whether their copy of Java’s API at the time could constitute fair use and thus not subject to copyright law is up to debate and my personal opinion doesn’t matter on it.
The Java Specification explicitly allows people to re-implement the Java Language as long as they follow a few guidelines (which google allegedly did not, hence the lawsuit) -- even if they decide to target a mobile platform.
Oracle in its case against Google is not arguing that "Java wasn't open-source at the time Google copied it". Oracle in its case against Google is not arguing that there was "a specific license explicitly disallowing mobile use". You on the other hand are arguing these things. That's where the problem lies: you're asserting infringement based on two fact claims that don't even match what Oracle's legal team presented to the courts.
(For that reason, your remark that "Everything else is irrelevant" is just bizarre and ironic—it's your comments here that are irrelevant... _None_ of the things you're saying are what the case is actually about.)
Here are some simple questions: to what extent does your knowledge of Oracle v. Google originate from secondary analysis and commentary about the case vs. direct knowledge (e.g. the briefs and testimony provided by Oracle and those who testified)? Do you have any firsthand experience reviewing the material that was presented in/to the courts? This is the problem with Internet peanut galleries. The answer to the last question can be solid "no", and yet commenters are undeterred from spewing nonsense from their gut that has no basis in reality.
But if you have done so, at least you as a presumably secondhand information source, could you give me a rebuttal on why am I wrong?
1. https://news.ycombinator.com/item?id=25847574
They also made a clusterfuck job copying Java APIs for their Android Java (aka Google's J++) dialect, where Java library authors have to hunt down what works exactly in each Android version.
They have only themselves to blame.
I'm super impressed by how the Truffle/GraalVM team has been able to turn this theoretical concept into a system that yields production grade compilers (although this Java on Java is not there yet).
"A very important detail of this implementation is that it’s implemented in Java. Java on Truffle is Java on Java! Self-hosting is the holy grail of Java virtual machine research and development." - now if you chuckled and imagined this as xkcd, I'm with you. Regardless I think this is really not some ultimate geek olympics fantasy, but that truffle and graal will be equally if not more transformative to PL reseaech than LLVM has been.
1: http://blog.sigfpe.com/2009/05/three-projections-of-doctor-f...
Self-hosting Java isn't a huge deal. IBM's Jalapeno/JikesRVM has been a JVM implemented almost entirely in Java for decades. It self-compiles by using the highest tier of its JIT on itself and then dumps the native code in memory to disk as an executable. Though, as far as I know, JikesRVM doesn't keep around its own bytecode (or an SSA representation of itself) for later dynamically re-optimizing itself or inlining parts of the JVM in user code hot spots. Hopefully a Truffle/Graal-based production JVM would be able to dynamically re-optimize/inline almost all of the JVM into hot spots in user code.
Currently Graal/Truffle implements the first Futamura projection: you give it an interpreter written in Java, it automatically generates a compiler (JIT). In this mode we have to keep the interpreter IR graphs (blueprints) around for partial evaluation.
With the first Futamura projection: Partial evaluator + interpreter + user code.
There's active work on implementing the second Futamura projection, where you partial evaluate the partial evaluator with respect to the interpreter, generating a specialized partial evaluator for that interpreter. With the second Futamura projection: (Specialized partial evaluator + interpreter) + user code.
This is truly fascinating and beautiful and the fact that it works for Java and not just a toy academic prototype is mind-blowing.
Okay, but if you're not bootstrapping, then you still need to keep OpenJDK around, and you're still hugely dependent upon big chunks of C++ code in OpenJDK, particularly in the garbage collector, unless they've added a crazy in-memory bootstrapping hook to OpenJDK to switch garbage collectors without shutting down the JVM.
Maybe there are some rapid development advantages to not requiring bootstrapping, but that means huge amounts of complexity around allowing re-definition of classes at runtime, particularly key JVM internals. Changing a tire without stopping the car is going to be very tricky.
- (Polyglot) scripting with Java
- Augmenting native images e.g. native javac with instant startup + annotation processors (very dynamic) running on Espresso
- A simple non-invasive JVM for constrained environments
- DCEVM-like features for developers
- Approachable academic playground
- Fast prototyping of JVM features e.g. it took our intern just two weeks to implement invokedynamic/MethodHandles
[1] https://github.com/hpi-swa/trufflesqueak/ [2] https://github.com/hpi-swa/polyglot-live-programming
On the other hand, it seems to me that these efforts and its resources would be better used developing a Webasm JVM implementation. Or at least kick-start the project.
We are missing a major platform in the JVM ecosystem and nobody seems to care.
I don't think graal atop wasm would be a particularly interesting use case. Wasm purports to be a portable, secure, performant application runtime. Graal purports to be a portable, secure , performant application runtime. The latter, however, has had many more engineer-hours thrown at its implementation, and is thus both more heavyweight (so it needs all the resources it can get) and more performant (so it makes sense to use it as the backing runtime).
I had thought the project had been abandoned, but I was wrong, it seems to be actively maintained and evolving on GitHub[1].
[0] http://teavm.org [1] https://github.com/konsoletyper/teavm/issues
The GC is written in Java, so if you have the translation to WASM that's covered as well.
As for objects, it's not clear to me how important it is to have runtime support for those.
It is basically made so c/cpp and other low level languages could be run inside browsers, and I don’t really see the point of it? Like, it is probably good for a video editor running inside a browser, but it will probably mostly entail all sort of mining, unnecessarily usage of cpu and the like. And it won’t magically improve performance, since at the boundary , calling javascript is really expensive, usually much more than writing the whole thing in js. Unless of course if someone does plenty of numerical computations.
Here is Flash doing the same with CrossBridge toolchain in 2011 with Unreal 3 citadel demo
https://adobe-flash.github.io/crossbridge/
https://www.youtube.com/watch?v=UQiUP2Hd60Y
Then there was PNaCl, but it wasn't good enough hence asmjs, and then WebAssembly, which is supposed to be secure.
Well,
"Usenix Security '20-Everything Old Is New Again: Binary Security of WebAssembly"
https://www.youtube.com/watch?v=glL__xjviro
TedDRA ANDF, IBM i, z/OS, CLR and BREW are all examples of bytecodes that support C and C++.
In a more ideal world, I would like to see the JVM underneath modern web browsers
Taking the argument to the extreme, assembly is a misfit.
For example there is also a Truffle implementation of C running on the JVM that runs some benchmarks faster than native C (due to runtime optimisations being better than static optimisations.)
I get what you are trying to say but if (!) you use JVM you are using an interpreter.
[1] https://web.archive.org/web/20190911160010/http://www.hanno....
You're using an interpreter to start with, but then it dynamically compiles machine code to run instead of using the interpreter.
But these terms are very fluid. I'd call HotSpot an interpreter, and a compiler, and both at the same time. It depends how you're looking at it at that moment.
Android had AOT compilation of Java bytecode (well, technically Dalvik bytecode) since 2013. It was upgraded to hybrid AOT/JIT mode in 2016. All of it was open-sourced, unlike Graal, where many key optimizations (such as auto-vectorization) remained proprietary.
Unlike GraalVM's AOT mode, Android VM can run existing Java libraries without any code changes — you don't have to painstakingly look for static variables and mark them for late loading, and reflection just works.
Most 3rd party commercial JDKs have offered it, specially for embedded deployment scenarios.
Besides Android Java isn't Java, rather a Google dialect with cherry picked packages from (nowadays) OpenJDK, most likely to die frozen in its current state as they embrace Kotlin über alles.
Android’s java is pretty terrible in comparison, just saying.