None of the IRs I’ve seen people try for Ruby does that. JSC’s DFG IR does a lot of this. Hence I don’t think folks have really tried the JSC approach for Ruby.
None of the IRs I’ve seen people try for Ruby does that. JSC’s DFG IR does a lot of this. Hence I don’t think folks have really tried the JSC approach for Ruby.
How do you deal with functions written in C++ which invoke functions written in JS?
Our C++->JS calling convention sucks, partly because we just avoid going down that path.
We do have a C++->JS call IC that we could use more.
But how would that work for Ruby, where C functions on the critical performance path are often third-party code that the compiler author has never seen before?
We'd need to let the JIT understand third-party unseen C functions. There are experiments to do that (Sulong, MIR, Rubinius sort of tried it) but I think it's more of an open problem than you're implying.
If you treat calls to unknown third-party C functions as an opaque native call then you're really going to struggle to build a meaningful compilation unit, in my experience.
"The blue parts show the new data-flow for MJIT. When building CRuby, we could generate MIR code for the standard Ruby methods written in C. We can load this MIR code as a MIR binary. This part could be done very quickly."
Flip thinks this isn't needed for Ruby - '[t]his project is trying to do too many things' - and that JavaScriptCore could do it already.
I think that MIR and TruffleRuby think they have to do something else (lifting C code into their IR) shows us that JavaScriptCore's approach isn't quite as immediately applicable as Filip thinks it is.
I buy that native code often calls back to Ruby.
I get why you would assume that therefore you need to make it fast for third party native code to call into Ruby. But that’s not how you want to think to succeed at VM optimizations.
You can make native code fast. It’s probably already about as fast as it’s going to be. Design a VM that makes it continue to be fast. Truffle won’t give you that since it will have to DBT the native code to make it fast (and that’s the good case).
I would bet you that first party native code that calls into Ruby is by far the most common kind of yield invocation. Like Array#each/map and equivalents for Hash. You want to treat those specially for two reasons:
- their fastest path for baseline code if done the way I describe is faster than any alternative. Baseline isn’t going to have a chance to inline arbitrary functions.
- they are likely to make up a large fraction of cases where native calls back to a Ruby.
For third parties, there’s a future where someone just exposes the JITing API that JSC gives to the DOM.
One major issue is the optimization scope is limited by how root traces are formed. This is fine in Lua but a much bigger issue for Ruby.
They try to infer specialization from an AST interpreter. That’s not on the same planet as what I’m talking about.
Do you mean an IR where all core library routines are first-class citizens, with their own nodes and information about their semantics encoded so that the compiler can reason about them?
With the caveat that core routines that are very cool don’t get included no matter how simple they are and warm/hot ones get opcodes even if it’s very annoying to do it.
Another big problem JavaScript doesn't have is that user C extensions are very often on the critical path in Ruby, so you must be able to optimise through them somehow. I don't think JavaScriptCore has any solutions for that we should be trying?
JSC has extensive support for fast calls into native code because of the DOM. For example we have the Snippet JIT that allows the DOM to turn hot functions into almost first class compiler ops. Like, callbacks in the DOM can dictate codegen and effect analysis.
Josh, the canonical JSC approach to the iterator callback problem wouldn’t be to use builtins. That wouldn’t necessarily achieve great perf. It’s certainly not great for the baseline tier. I think the only good option is to make those iterators (like Array#each) be intrinsic as fuck: all call ICs can ask for inline machine code generation of the loop along with the call back to the passed in block. You could imagine this enabling inlining of the loop and its body even in the baseline JIT.
I get that people probably see the builtins in JSC and start having wild fantasies about what this can achieve. In reality for things where perf matters, you deploy ICs and custom template codegen.
I don't really see a good solution here other than lifting the control flow to Ruby with specific bytecode ops to minimize overhead and a good, typed FFI like LuaJIT has? Am I missing something?
Other than that just try to make calls into JITed code as fast as possible. But that needs to be the backup plan, with the main plan being that perf-sensitive extensions play with the JIT.
And by the way - there isn't really any proper interface at all! There's just basically the whole internals of the C Ruby implementation exposed to C extension authors. There's no abstraction with handles and things like that.
There was an effort at one point to get C extension authors to use the FFI instead, and a JNI-style API to permit a moving garbage collector was proposed, but these approaches weren't successful in gaining any momentum.
All in all... that's why we don't just do it the JavaScriptCore way. We have different constraints to yours.
The thing about moving GC is a red herring. Ruby wouldn’t benefit from it. JSC doesn’t use moving GC.