NativeJIT – JIT compilation of expressions involving C data structures
github.com
github.com
Interestingly, it requires the functions called by the JIT'd code to be side-effect free, since it guarantees it will call any given function invocation at least once, since it evaluates both sides of any but top level branches. See "Design Notes and Warnings" in https://github.com/BitFunnel/NativeJIT/blob/master/Documenta...
On top of that, LLVM is a massive dependency.
I only looked at NativeJIT a few minutes but it seems to solve both of these issues. It's pretty small in terms of API surface // total amount of code and is explicitly designed for a usecase where compilation happens as part of a user-issued query (that needs to be fast).
That's arguable. Ravi[1] is a Lua implemented with an LLVM-based JIT. The benchmarks here[2] compare runtime performance of Ravi vs. LuaJIT. The Ravi timings don't include compilation time, the LuaJIT timings do. Sometimes LuaJIT siginficantly outperforms Ravi, sometimes it's on par, sometimes worse. But the results don't suggest that there is less (effective) optimisation going on.
[1] https://github.com/dibyendumajumdar/ravi
[2] http://the-ravi-programming-language.readthedocs.io/en/lates...
Optimization pipelines are invariably language-specific. If you were to somehow make LuaJIT into a C/C++ compiler without doing any work on its optimizer, you would get similarly poor results.
Clang -> Sulong -> Java bytecode -> Luje [1]? May need more turtles :)
Technically, one reason -of several- this probably wouldn't work without significant effort is that Sulong relies on Graal's FFI intrinsics. But those should be convertible to LuaJIT FFI calls (in theory).
You can't really compare LuaJIT and LLVM directly, because LuaJIT is a trace compiler. A LuaJIT trace is a linear sequence of instructions with no control flow, which makes optimization much easier. LLVM on the other hand is statically compiling code consisting of many basic blocks and potentially complicated control flow between them.
So LLVM's job is much harder. LuaJIT's optimizer doesn't have to work nearly as hard to get a similar quality of code.
Will seriously consider using this to speed up expression execution in EventQL [0] (we shied away from llvm so far because it's such a massive dependency).
For example, if you evaluate sum(sqrt(x^2 + y^2) * 3) in Renjin, and x or y happen to be very long vectors, then we'll jit out a JVM class for this specific expression that would look something like this in Java:
class JittedComputation1E5374A3 {
SEXP compute(SEXP[] args) {
double[] x = args[0].toDoubleArrayUnsafe();
double[] y = args[1].toDoubleArrayUnsafe();
double sum = 0;
for(int i=0;i<x.length;++i) {
double xi = x[i];
double yi = y[i]
sum += Math.sqrt(xi*xi+yi*yi) * 3;
}
return DoubleVector.valueOf(sum);
}
}
The computation is specialized to the types of x and y, so if for example x is a sequence 1:1000000 then a new class gets written for that doesn't even use an array for x.The speedup is so impressive that even if you don't cache the compiled expression you see dramatic improvements: http://www.renjin.org/blog/2015-06-28-renjin-at-rsummit-2015...
You basically really want this in any case where you're applying a dynamic expression to a large dataset.
We use the technique in Renjin, an R interpreter built on the JVM, for vectorized expressions such as sqrt(x^2+y^2) where x and y are long arrays.
You can implement this with function pointers, or interfaces in Java, but if you're evaluating the expression millions of times, then the cost of the indirection is huge, and worse yet, it gets in the way of the processor's branch prediction and pipelining.
If you can compile the same dynamic expression down to straightline machine code, the impact is pretty fantastic: http://www.renjin.org/blog/2015-06-28-renjin-at-rsummit-2015...
Ex:
// nativeJIT DSL
auto & area = expression.Mul(rsquared, expression.Immediate(PI));
auto function = expression.Compile(area);
// C++
auto area = rsquared * PI;Out of interest, can you recommend any others that are actively maintained?