Deep Dive into PHP 8's JIT
thephp.website
thephp.website
https://github.com/php/php-src/pull/5874
https://github.com/php/php-src/commit/4bf2d09edeb14467ba7955...
AFAIK tracing JITs are generally inferior to method-based ones, which is why none of the major JavaScript engines use tracing. Their only advantage seems to be (relative) simplicity, which is essential for the lightweight LuaJIT but not for PHP.
(I don't know if the HotSpot folks ever considered adopting the changes.)
When compiling at method-granularity, with incomplete information about types, complex features are required e.g. On-Stack Replacement (if the current method becomes "hot", it needs to be replaced by a compiled version), Inline Caches, Hidden Classes and Deoptimization.
In a trace-based JIT, all these features can be replaced by "just compile another trace".
LuaJIT does without hidden classes or property caching because on-trace it's able to eliminate table access and allocation. Performance tanks on OO Lua once you fall off trace. You can work around it with "compile another trace" in a sense that works as long as you have infinite cache.
There's a lot more to it than just compiling what the interpreter does. CRuby's JIT has a different approach using C templating rather than Dynasm and templating/context threading but it also fails to improve performance much for the same reasons.
Tracing JITs are really good at optimizing a small scope. What a tracing JIT does with recording a single path of linear control flow the method JITs also try to do with branch profiling and pruning.
It works really well for a particular style of Lua. For a language like Ruby or PHP it's a lot harder to make a tracing JIT work well on existing code.
The main problem is tail duplication causes the number of traces to increase exponentially with each branch. You have to either have trace heuristics which keep the traces very short or go half way to a method JIT and build a control flow graph.
I got to the tail explosion problem building a tracing JIT for CRuby and gave up.
But anyway, kudos to the PHP folks for yet another improvement. It’s been a long time since I used it but I’m continually impressed by how far it’s come. And it’s a useful test to how flexible a developer is willing to be to ask them to code some PHP, there are fewer and fewer reasons against it these days other than personal style preference.
Cpython is interpreted, pypy has JIT, mypyc is compiled. It's the implementation, or even implementation's specific runtime options that decide about things like that, not the language.
C can be interpreted too: https://en.wikipedia.org/wiki/CINT
Compilation is just interpreting what a program does and producing machine code (or other code in the case of transpilers) that computes the same thing.
Interpreters are just that, but instead of producing code, they run code in themselves thats computes the results directly.
One could even consider machine code just “obfuscated” assembly code. In that sense, machine code is just another language that also could be interpreted (intepretation-based emulators like Bochs) or compiled again (JIT-based emulators like DOSBox).
This brings up a slightly related question: is there a language that can’t be compiled?
There is some question whether Perl could be, because parsing it without running it has some ambiguity. https://www.perlmonks.org/?node_id=663393
[0] https://en.wikipedia.org/wiki/Partial_evaluation#Futamura_pr...
[1] http://blog.sigfpe.com/2009/05/three-projections-of-doctor-f...
Not sure if you're asking this platonically, but Futamura shows us that if you can built an interpreter then you can always transform that automatically to be a compiler.
This is used in practice by some compilers today - they automatically produce a compiler from an interpreter!
People like speaking in the general case. It saves time from enumeating any inconsequential / statistically not relevant exception.
The real problem is people misunderstanding casual discussion absolutes (which mean "for the large majority/for the ones people care about") with mathematical absolutes (i.e. "X is Y for each and every X")
I didn't explicitly say that JS doesn't come with JIT, because I'd be putting Node.JS and every browser engine in the same bag. I simply can't be sure about all of them.
I'll add very soon this caveat that the dialect (js, php, python) doesn't really matter and make it clear that I'm talking about engines, possibly directly point to them.
I hope you understand, though, that I wanted to make JIT as a concept understandable. The language comparisons are secondary.
Thanks a lot for your feedback!
Truth can be harsh, but people sometime overvalue to such an extent is hard to understand.
sum = 0
for i in xrange(n):
for j in xrange(i):
sum += A[i][j]
And when they reach that milestone, some people call it "as fast as C".Never mind that that's not what people actually write in Python or PHP. It's a synthetic benchmark, not a real workload.
The workloads in those languages are generally oriented around strings, hash tables, and function/method calls.
And the JITs don't seem to do nearly as good a job there. I tested PyPy on Oil [1] a few years ago, and it made it slower, not faster. And it used more memory. (Though PyPy is an amazing project in many respects.)
You usually don't care how your matrix multiplication/regex matching/unicode normalization/JSON parsing is implemented, but people had to make those, and they are users of the language too.
Even though it might not change the bottom-line for your high-level app.
Julia is a dynamic language that seems to do better because it was designed for this purpose.
But it doesn't seem to have panned out in practice in Python, or PHP as far as I know. Those languages have huge piles of C, and whenever you call into C, the JIT gets confused. People don't seem to rewrite their huge piles of C in Python or PHP. In Python, it's more likely Cython.
I'd like to see pointers to counterexamples -- where people actually wrote some C-like code in Python or PHP and let the JIT do its work. I haven't seen it, aside from the PyPy project itself, and maybe a few other examples. I think you would still take a significant performance hit.
The issue is that C compilers in 2020 are even better at compiling the example I showed. They do amazing things with that kind of code that state-of-the-art JITs don't in practice.
So if a Python JIT does 10-50x better than CPython on a numeric workload, that sounds impressive, but it's still slow compared to C.
And again they don't get 10-50x on string/hash/method call workloads. I think they're lucky to get 2x in some of those cases.
Well, C hardly is, either...
However a call to a shared library that isn’t linked against your language API is not very expensive as you have a much better handle on the values that are escaping and can make much better optimization choices.
In the Truffle project we are using an LLVM bitcode interpreter that allows us to JIT right through that language boundary and still link to native shared libraries. This means people shouldn’t have to rewrite their C extensions and we can hopefully still run the combination of high level language and C extension faster.
JSC and LuaJIT have simpler ways to deal with calling native code which might do weird stuff.
I think a workload like this is more common in the PHP world. Not saying that others don’t exist, but handling routing, queries, cached content is very different from simply doing mathematical/memory intensive applications.
JIT compilers have different optimizations available like:
Fast inlined heap allocation (normally much faster than malloc/free). V8 even does allocation combining
Transparent ropes for strings
High level alias analysis for hash tables and objects
Inlining dynamic dispatched functions and dynamically loaded functions
Any language can eventually reach that point with enough money, time, and in C's case doing 200+ optimizations with unexpected results.
So there's nothing really funny about it....
Since 25 years, time to learn that implementations and languages are not the same.
Could you please point an implementation detail where a JIT-capable engine doesn't include interpretation in its runtime?
In every case, thanks a lot for your feedback!
It's not super exotic to do this.
For example in .NET, MSIL goes directly into a pipeline that produces native. You can easily validate that RyuJIT has no interpretation.
Or for example, watchOS applications packaged with bitcode, get JIT compiled at installation time.
- https://phabricator.wikimedia.org/T176370
- https://phabricator.wikimedia.org/T229792
- https://launchdarkly.com/blog/how-the-wikimedia-foundation-s...
But this Blog, a single page has 5 Google Ads in it. I dont mind one or two, top and bottom. But 5, right in the middle of every section.
Is that not a thing anymore? Genuinely curious
Edit: Yes, some initial searches reveal that it's not a thing anymore.
But if you believe the amount is so harmful for your experience, don't worry. I'll be more than glad to reduce this amount.
Cheers!
Cheers!
Deep dives were handled before: https://wiki.php.net/rfc/jit And its discussion https://externals.io/message/103903