Mir: A lightweight JIT compiler project
developers.redhat.com
developers.redhat.com
It seems like Ruby needs a profile guided optimizer, which means building an IR that is suitable for profile-guided optimization. That’s way different from classic IRs like this since it means having provisions for OSR exit.
I recommend looking at these slides to learn how to do it.
http://www.filpizlo.com/slides/pizlo-speculation-in-jsc-slid... http://www.filpizlo.com/slides/pizlo-splash2018-jsc-compiler...
However, the reason that this is not so simple is that Ruby is a vastly more complicated language than JavaScript (I've worked on implementing both.) Ruby has an enormous standard library and most Ruby programs are just endless calls to the that library, so your compiler must be able to understand the library semantically, either by rewriting it in Ruby (not likely at scale in MRI) or by adding tens of thousands of individually optimised intrinsics (again not likely).
Ruby is digging itself further into a local optima with these approaches optimisations, rather than looking further around for a better global optima.
Maybe the reason why Ruby optimization has problems is nobody has done it the JSC way.
Yes it's the same problem and it's unique to neither Ruby nor JS - but the problem is just scale.
Ruby has a larger library, so needs more rewritten from the current C into Ruby+builtins. And then that rewrite doesn't maintain Ruby semantics (the C API currently used does not exactly match Ruby semantics) so a rewrite matching semantics is often very hard.
> Maybe the reason why Ruby optimization has problems is nobody has done it the JSC way.
People have been trying the approach you use in Core (for over a decade, starting with Rubinius, now TruffleRuby and others), but I think (based on practical experience working on optimising both languages) that it's just a larger problem in Ruby which is why it hasn't been conquered yet.
None of the IRs I’ve seen people try for Ruby does that. JSC’s DFG IR does a lot of this. Hence I don’t think folks have really tried the JSC approach for Ruby.
They try to infer specialization from an AST interpreter. That’s not on the same planet as what I’m talking about.
Do you mean an IR where all core library routines are first-class citizens, with their own nodes and information about their semantics encoded so that the compiler can reason about them?
With the caveat that core routines that are very cool don’t get included no matter how simple they are and warm/hot ones get opcodes even if it’s very annoying to do it.
Another big problem JavaScript doesn't have is that user C extensions are very often on the critical path in Ruby, so you must be able to optimise through them somehow. I don't think JavaScriptCore has any solutions for that we should be trying?
JSC has extensive support for fast calls into native code because of the DOM. For example we have the Snippet JIT that allows the DOM to turn hot functions into almost first class compiler ops. Like, callbacks in the DOM can dictate codegen and effect analysis.
Josh, the canonical JSC approach to the iterator callback problem wouldn’t be to use builtins. That wouldn’t necessarily achieve great perf. It’s certainly not great for the baseline tier. I think the only good option is to make those iterators (like Array#each) be intrinsic as fuck: all call ICs can ask for inline machine code generation of the loop along with the call back to the passed in block. You could imagine this enabling inlining of the loop and its body even in the baseline JIT.
I get that people probably see the builtins in JSC and start having wild fantasies about what this can achieve. In reality for things where perf matters, you deploy ICs and custom template codegen.
I don't really see a good solution here other than lifting the control flow to Ruby with specific bytecode ops to minimize overhead and a good, typed FFI like LuaJIT has? Am I missing something?
Other than that just try to make calls into JITed code as fast as possible. But that needs to be the backup plan, with the main plan being that perf-sensitive extensions play with the JIT.
And by the way - there isn't really any proper interface at all! There's just basically the whole internals of the C Ruby implementation exposed to C extension authors. There's no abstraction with handles and things like that.
There was an effort at one point to get C extension authors to use the FFI instead, and a JNI-style API to permit a moving garbage collector was proposed, but these approaches weren't successful in gaining any momentum.
All in all... that's why we don't just do it the JavaScriptCore way. We have different constraints to yours.
The thing about moving GC is a red herring. Ruby wouldn’t benefit from it. JSC doesn’t use moving GC.
How do you deal with functions written in C++ which invoke functions written in JS?
Our C++->JS calling convention sucks, partly because we just avoid going down that path.
We do have a C++->JS call IC that we could use more.
But how would that work for Ruby, where C functions on the critical performance path are often third-party code that the compiler author has never seen before?
We'd need to let the JIT understand third-party unseen C functions. There are experiments to do that (Sulong, MIR, Rubinius sort of tried it) but I think it's more of an open problem than you're implying.
If you treat calls to unknown third-party C functions as an opaque native call then you're really going to struggle to build a meaningful compilation unit, in my experience.
"The blue parts show the new data-flow for MJIT. When building CRuby, we could generate MIR code for the standard Ruby methods written in C. We can load this MIR code as a MIR binary. This part could be done very quickly."
Flip thinks this isn't needed for Ruby - '[t]his project is trying to do too many things' - and that JavaScriptCore could do it already.
I think that MIR and TruffleRuby think they have to do something else (lifting C code into their IR) shows us that JavaScriptCore's approach isn't quite as immediately applicable as Filip thinks it is.
I buy that native code often calls back to Ruby.
I get why you would assume that therefore you need to make it fast for third party native code to call into Ruby. But that’s not how you want to think to succeed at VM optimizations.
You can make native code fast. It’s probably already about as fast as it’s going to be. Design a VM that makes it continue to be fast. Truffle won’t give you that since it will have to DBT the native code to make it fast (and that’s the good case).
I would bet you that first party native code that calls into Ruby is by far the most common kind of yield invocation. Like Array#each/map and equivalents for Hash. You want to treat those specially for two reasons:
- their fastest path for baseline code if done the way I describe is faster than any alternative. Baseline isn’t going to have a chance to inline arbitrary functions.
- they are likely to make up a large fraction of cases where native calls back to a Ruby.
For third parties, there’s a future where someone just exposes the JITing API that JSC gives to the DOM.
One major issue is the optimization scope is limited by how root traces are formed. This is fine in Lua but a much bigger issue for Ruby.
The key is:
- large opcode set and an architecture that tries to amortize the pain of lots of opcodes.
- excellent support for speculation and effects analysis.
(for the uninitiated, I kid. Sample of previous naming controversy at https://news.ycombinator.com/item?id=14309903)
"Plans to try MIR light-weight JIT first for CRuby or/and MRuby implementation"
"MIR is strongly typed"
Is there an explanation of how the project bridges the gap between dynamically-typed Ruby and statically-typed MIR?More generally, I'd love to see something like MRuby+MIR be successful. It would be great to see an alternative to the aging LuaJIT.
Dynamic types and garbage collection would then be implemented on/for the abstract MIR machine.
My understanding is that MIR doesn’t tackle this problem, so the code does stay pretty dynamic when compiled, just like its sister project YARV MJIT.
They both need an intermediate profiling mode to specialise and monomorphise but nobody is building that as far as I know.
<quote> No SSA (single static assignment form) for:
Faster optimizations for short optimizations pipeline and small functions (a target usage scenario) Currently SSA could be used only for two optimizations (CCP and GCSE). SSA usage would mean 4 additional passes over IR. If we implement more optimizations, SSA transition is possible when additional time for expensive in/out SSA passes will be less than additional time for non-SSA optimization implementation Simpler and more compact generator code because we can avoid to implement a lot of nontrivial code (for dominator and dominator frontier calculation, a good out of SSA code) </quote>
Of course it's debatable what "modern" and "non trivial" mean, and if they apply to this project. It's very small, but on the (one!) small benchmark the author cites it seems to do quite well, so it's certainly not completely naive.
For whatever it's worth, CompCert doesn't use SSA either, and that's certainly a non-trivial compiler, though arguably the non-triviality does not stem from any advanced optimizations it does.
CRuby with MJIT
MIR ( This )
JRuby ( Ruby on JVM )
TruffleRuby ( Ruby on Graal )
Artichoke ( Ruby on Rust )
And I remember someone mentioned making Ruby with Tracing JIT. ( Not Topaz ) Unfortunately My Google fu is not good enough I can no longer find it.
https://llvm.org/devmtg/2009-10/Phoenix_AcceleratingRuby.pdf
EngineYard took Rubinius in a few directions, but I think the main lasting impact was all the RSpec work they did along the way.
Ideally it handles register allocations and generated function caching as well. The current JIT assembler (ASMJIT, Xbyak) requires you to handle register allocations. LLVM is as mentioned quite a heavy dependency to have.
The article mentions Cranelift, but that's a 'middleweight JIT' with a proper SSA IR. I was surprised to see LibJIT has more LOC than Cranelift - I thought it was lighter. (Imperfect proxy for runtime 'weight', of course.)
If you want a lightweight portable JIT engine, there's already GNU Lightning [0], and the atrociously-named Lightening fork [1] (used in the new JIT in the GNU Guile Scheme interpreter, which turned up on the HN front page recently).
Here's a 1996 paper (preprint) on a research JIT named VCODE which executed around 8 instructions to generate each instruction in its output. [2] (Sadly it was never released, as far as I can tell, and is presumably long dead.)
Anyway, with all that said, I wish this project well. No-one's managed to get good performance out of Ruby yet, so it's certainly ambitious. Google gave up on Unladen Swallow, and that was a JIT for Python, which, as I understand it, is more amenable to JIT than Ruby. Even failing that, having a quality rival to GNU Lightning would be worthwhile.
[0] https://www.gnu.org/software/lightning/
[1] https://www.wingolog.org/archives/2019/05/24/lightening-run-...
[2] http://www-leland.stanford.edu/class/cs343/resources/vcode-a...
Anything you write in MIR can be compiled by MIR. Parrot only JITs Parrot bytecode, so the applications are much more limited.
That's the point where "C is nice and simple, it's easy to whip up a compiler" invariably turns into "why the #^(&)# didn't I use an existing frontend?". Real-world C code is messy. The entire sub-project of implementing a C compiler is a needless distraction that will turn into a huge time-suck with zero benefit to the author.
See also: "Why Says C is Simple?" https://people.eecs.berkeley.edu/~necula/cil/cil016.html
Ruby was my go to language for a long dry spell when I had little Lisp development jobs (except for Clojure). Ruby is a great language, as Matz says, Ruby is designed for developer happiness. I stopped using Ruby when more Common Lisp work came my way and then I used Python for five years of deep learning work. I have favorite languages but I used what customers wanted.
That said, I still keep up with Ruby news.
I feel like this list is missing some more, but yeah, there's been a few 'Mir' projects.
It's so terribly confusing because LLVM itself defines a .mir (Machine IR) that is a syntax between its IR and the backend (which could be thought of as mid-level IR). When I heard about Rust's MIR, I assumed it was some clever way to generate target-dependent IR.
The good news is that we now have M(L)IR: the one IR to bring them all and in the darkness bind them.