When FFI function calls beat native C
nullprogram.com
nullprogram.com
I understand calling dlopen() and dlsym()
And I understand this idea of a PLT and it's indirection
But this idea of something external to the JIT'ed program being JIT'ed I do not understand.
Does it mean it inlined the instructions of the external function into the JIT'ed code?
What's being JIT'd is the Lua code into machine code, and where Lua would need to call a C function, it (apparently) is just emitting a `call the_external_fn` into the JIT's resulting assembly. That's a direct function call, so it's about as fast as you're going to get, but somewhat counter-intuitively, it'll be faster than C, as we don't do anything with the PLT or indirection. Just call the function.
aka, upon DLL invocation (on Win32 they call it DllMain I guess, I've seen it call ctor/dtor on *nix), spawn the runtime, and expose FFI functions?
http://www.drewtech.com/support/passthru.html
This spec is really big on FFI exposing functions. I always find it an edge case when trying to play with certain technologies (like the one in the article).
LuaJIT also allows you to pass a Lua function as a callback to a C function expecting one, subject to some limitations.
I'm not sure if either of those answers your question though.
Something that might be relevant is that that LuaJIT can JIT-optimize FFI calls from Lua but can't/doesn't optimise calls into Lua made via C.
I might be able to go digging for the reference (at this point figure it's best to just reply for now, not sure if you'll see this) but I've read that "the approach" recommended to solve this problem is to move the main `for(;;)` / `while(1)` loop into Lua and have LuaJIT repeatedly FFI-call C, because that's the path that can go the fastest.
With modern CPU architectures that can prefetch/speculate over indirect jump/call making it comparatively expensive to indirect call in a tight loop (the idea behind PLT is that in contrast to indirect call through GOT it should more clearly signal the intent to the CPU). On the other hand the tight loop is certainly important and this effect will not be so pronounced on any kind of practical code (because of the BTB pressure).
Also, this is interesting observation for various discussions about overhead of late binding ("virtual" in C++) as similar overhead is already there for almost any cross-object function call in PIC dynamic binary.
(Linked from the article)
So this is a Linux or *nix specific quirk, rather than a C quirk. Apologies if my memory isn’t accurate.
On typical ELF platform with dynamic linking you get PIC compiled binaries with the added overhead of GOT and PLT. One thing to keep in mind is that on many traditional unix RISC platforms there is something similar to GOT and PLT even in statically linked binaries (no word-sized immediaties on RISC platforms…).
https://www.pypy.org/posts/2011/02/pypy-faster-than-c-on-car... - JIT'ing across compilation units
https://www.pypy.org/posts/2011/08/pypy-is-faster-than-c-aga... - JIT'ing % interpolation.
(Wow, those are 11 years old. I remember when PyPy was a new project.)
The C code contains no optimization annotations either, the compiler could be inlining the indirect benchmark and/or devirtualizing the indirect call itself.
Can you cite where in the article it addresses the fact that the assembly snippet is not an apples to apples comparison with the C code?
> If you have a specific criticism to make
Pointing out that the assembly is not an apples to apples comparison with the C code is a specific criticism.
Well, you can do
MOV rax, 0x1122334455667788
PUSH rax
RET
in this case. Still direct, just a bit slower. Wonder if modern CPUs speculate past this construction.> The downside to this approach is slower loading, larger binaries, and less sharing of code pages between different processes. It’s slower loading because every dynamic call site needs to be patched before the program can begin execution. The binary is larger because each of these call sites needs an entry in the relocation table. And the lack of sharing is due to the code pages being modified.
It probably could be improved by essentially doing the same thing that luajit is doing and inventing a JIT code loading mechanism, but that would be hard to get a lot of buy-in. You need to change the ELF standard and get the likes of GCC and LLVM on board with this new paradigm.
gcc has an option to reduce the overhead with no-plt (https://gcc.gnu.org/onlinedocs/gcc/Code-Gen-Options.html). so on x86_64 it just does call[rip + got_offset]. this is still an indirect call but it reduces the number of calls by 1.