In libffi you built up descriptor objects for functions. These are run-time data structures which indicate the arguments and return value types.
When making a FFI call, you must pass in an array of pointers to the values you want to pass, and the descriptor.
Inside libffi there is likely a loop which walks the loop of values, while at the same time traversing the descriptor, and places those values onto the stack in the right way according to the type indicated in the descriptor. When the function is done, it then pulls out the return according to its type. It's probably switching on type for all these pieces.
Even if the libffi call mechanism were JITted, the preparation of the argument array for it would still be slow. It's less direct than a FFI jit that directly accesses the arguments without going through an intermediate array.
FFI JIT code will directly take the argument values, convert them from the Ruby (or whatever) type to the C type, and stick it into the right place on the stack or register, and do that with inline code for each value. Then call the function, and convert the return value to the Ruby type. Basically as if you wrote extension code by hand:
// Pseudo-code
RubyValue *__Generated_Jit_Code(RubyValue *arg1, RubyValue *arg2)
{
return CStringToRuby(__targetCFunction(RubyToCString(arg1), RubyToCInt(arg2));
}
If there is type inference, the conversion code can skip type checks. If we have assurance that arg1 is a Ruby string, we can use an unsafe, faster version of the RubyToCString function.The JIT code doesn't have to reflect over anything other than at worst the Ruby types. It doesn't have to have any array or list related to the arguments. It knows which C types are being converted to and form, and that is hard-coded: there is no data structure describing the C side that has to be walked at run0-time.
Yes it's probably worse than doing the jit in Ruby interpreter, since there you can also inline the type conversion calls, but there principles are the same.
Edit: This is wrong, see comments below.
Running in a debugger an ffi_call to a "int add(int a, int b)" leads me to https://github.com/libffi/libffi/blob/1716f81e9a115d34042950... as the assembly directly before the function is invoked, and, besides clearly not being JITted from me being able to link to it, it is clearly inefficient and unnecessarily general for the given call, loading 7 arguments instead of just the two necessary ones.
(the tramp.c linked in a sibling comment is for "reverse-FFI", i.e. exposing some dynamic custom operation as a function pointer; and its JITting there amounts to a total of 3 instructions to call into precompiled code)