JIT and Ruby's MJIT
engineering.appfolio.com
engineering.appfolio.com
Much of the problem has to do with C extensions, despite the team's effort to even have C Code Compiled with TruffleRuby, they also had to go all the way to fix all the C library that didn't work with TruffleRuby due to undefined behaviours.
Graal has an open source license, and all corporations are alike, regardless of what hippie developers think, showing to the man and such stuff.
Instead other language stacks will get adopted into detriment of Ruby, e.g. Go.
lol.
> "what you think of Oracle is even truer than you think it is. There has been no entity in human history with less complexity or nuance to it than Oracle"
> "this company is about one man and his alter ego and what he wants to inflict upon humanity"
> You need to think of Larry Ellison the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower hates me' -- lawnmower doesn't give a shit about you, lawnmower can't hate you. Don't anthropomorphize the lawnmower. Don't fall into that trap about Oracle.
https://www.youtube.com/watch?time_continue=2318&v=-zRN7XLCR...
> That’s a pretty significant footnote.
Oops.
> In addition to memory usage, there’s warmup time. With JIT, the interpreter has to recognize that a method is called a lot and then take time to compile it. That means there’s a delay between when the program starts and when it gets to full speed.
Why not cache the compiled methods to make warmup a once-per-version delay? It would be a JIT/precompiled hybrid. Call it gradually compiled.
>> That’s a pretty significant footnote.
> Oops.
FWIW / from what I remember, generating high-performance native code for Rails was hard (or at least significantly different from optimizing standard benchmark type code). The OMR JIT for Ruby had similar challenges, with at least a few 'at least we didn't make it worse' moments.
I was only tangentially involved, but I vaguely recall Rails needing (or at least appearing to need) a lot more inlining and string manipulation optimizations to fly.
That's what Bootsnap does, and it now comes standard with Rails. https://github.com/Shopify/bootsnap
You seem to be suggesting that Smalltalk implementations are a whole lot faster than current Ruby ?
They are a whole lot faster than Matz's 2008 Ruby 1.8.7 but so is current Ruby:
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
and
"Largest Provider of Commercial Smalltalk
...Cincom is the largest commercial provider of Smalltalk in the world, with twice as many partners and customers than all other commercial providers combined."
"Generally they are in the same league."
>When a method has been called a certain number of times (10,000 times in current prerelease Ruby 2.7), MJIT will mark it to be compiled into native code and put it on a “to compile” queue. MJIT’s background thread will pull methods from the queue and compile them one at a time into native code.
I would have thought that you wouldn't wait for 10,000 iterations but just start from the beginning and compile the methods with the most calls and keep the compiler thread busy. Flush less frequently or no-longer needed compiled calls and limit the total to some defined resource usage cap. You'd probably win overall.
https://github.com/dotnet/coreclr/blob/master/Documentation/...
30 seems crazy low to me. That seems like you'd be spending a bunch of compute time early on compiling stuff that may be only used during the start-up stages of your code.
I haven't kept up with Ruby JITs, so I don't know if / how mjit handles inlining, but what I remember from my J9 / OMR days is that choosing what to inline made a massive difference in compiled code performance. Inlining too much was a great way to hobble performance.
> Flush less frequently or no-longer needed compiled calls...
The technical complexity of this process should not be underestimated. Given the minimalism of mjit's approach, I could easily see such a strategy being unviable without infrastructure investment that would (at least appear to) be significantly further along the effort / reward curve.
That said, there are plenty of Rails apps that do monkey patch Rails internals, for better or worse.
Most JavaScript JITs?
First of all there are many JVM implementations, the commercial ones used to offer AOT compilation as well, going back to early 2000.
Then the ones that have JITs offer multiple flavours.
One way is to initially interpret the bytecodes, after enough information it is gathered, the first level JIT gets into action and compiles that block into native code, here block is usually a function, but can be something else.
This first level compiler is rather simple and does only basic optimizations.
The application keeps profiling execution and eventually notices that the already compiled block (into native code) keeps being used significantly, now it is time to bring the big brother JIT, which is somehow equivalent to -O3 on gcc, and recompile again to native code using all major optimizations.
Other JVMs (like JRockit) never interpret, when they start the first level is already the dumb level one compiler to native code.
Then all of them now support JIT caches, meaning after a run, the JITed methods get saved and re-used by next execution, so the profiler gets to learn from previous runs, and execution of the system already starts from a much better performance state.
It initially interprets, and when a specific threshold is reached (you can configure it), the C1 compiler gets called into action doing basic optimizations.
After awhile if that native generated code keeps getting even more hot, the C2 compiler (the one with -O3 capabilities) gets called into action.
In both cases the optimized code gets safety guards to validate that the assumptions made by the JIT are still valid. For example if a dynamic dispatch always lands on the same method, then it gets replaced by a direct call instead. Even that is proven wrong, then the JIT throws the optimized code away and starts with the new assumptions.
Then in what concerns OpenJDK, you have actually 2 C2 JIT compilers available, HotSpot written in C++ and still the default, and Graal written in Java taken from GraalVM (nee MaximeVM) project. Currently Graal is much better than HotSpot in escape analysis for example, but worse in other scenarios.
In both cases, OpenJDK has inherited the JIT cache infrastructure from JRockit, so you also get to save the native code between runs, and start much faster in consequent runs.
As note, even though it is usually not a good idea, if you set the interpreter threshold to zero, then C1 kicks right at the beginning, but it won't have any information available, so the generated code is going to be most likely worse than just interpreting.
If you actually need a compiled Ruby binary mRuby is the best option but it has its own drawbacks of course.
It may work a bit better now, but it's still very much not tuned for that. The result would probably still be slower than not JITting, for the same reasons mentioned for 2.6 in that link.