We get first-class cross-language FFI to truffled language implementations: Python, Ruby, R, Javascript with shared objects, which can JIT together. The idea is that fully fleshed out this would give Haskell a clean path to reach over to pytorch to run an AI model, take the answers, smash them through some statistical analysis using an obscure R library and send the result to d3.js to get pretty visualizations. Are we there yet? Hell no.
We get access to the java incubator vector API allowing us to JIT SIMD kernels for the target platform with less pain. Haskell has just been bad at high speed SIMD code since I found my way into the ecosystem, and it hasn't shown much sign of getting better. This lets me theoretically sidestep those issues.
Using Sulong for C/C++ FFI means we don't give up native cbits and host code, unlike old bad hard-to-use JVM language ports like JPython.
Future work could for instance make Natural number code use java deoptimization paths keeping it generally in unboxed ints until you finally need something too big through a given codepath. Again, just more paths to potentially make things faster or easier to use.
Without Truffle/Graal the limitations of the JVM are just too severe. It is an awkward runtime full of limitations. The lack of proper tail call optimization for instance would kill GHC-style evaluation performance, and has left a long list of functional language corpses in its wake that tried to make the JVM their home.
With the "Cadenza trick" I use to eliminate that we can have highly performant loops. Using assumptions lets it compile with the equivalent of GHC's single threaded runtime and then downgrade performance to the equivalent of the multi-threaded runtime when you first call a threading operation, so users don't have to pick, they just get the benefits. I, in turn, get headaches.
We get stuff that GHC just will never even try to get around to. CompressedOOPS give us 32 bit pointers on 64 bit platforms if the heap is < ~32GB. So about twice as much stuff can fit into cache especially in a language with as many pointers running around as ours compared to the normal Haskell runtime.
My goal is to keep pushing forward with the bits that this can do that GHC can't do, and to generally try to get to or maintain parity wherever possible for the things that GHC can do.
It also helps us tease apart where the line should be for GHC as a runtime system vs. GHC as a compiler. The former has been languishing behind comparatively, ever since Simon Marlow left GHC headquarters to go fight spam for Meta.
I wrote a blog post targeted at an insider audience to tease at something that was coming. It has since spread outside of that audience. There's no gatekeeping going on here, no secret enlightenment required, just a poorly drafted research note thrown over the fence to his peers to see if anybody else might be interested in his new favorite toy by a very busy person. I'm sorry.