PyPy: A Faster Python Implementation
pypy.org
pypy.org
Is that accurate? What are the pros & cons of this approach compared to a normal tracing JIT, e.g. LuaJIT?
If you have a program and represent it as a series of VM instructions, and you try to jit them individually the kit compiler has very limited ability for optimization. Plus, after each such instruction the execution comes back to the piece of the interpreter that selects the next instruction to execute: essentially a giant switch statement that makes branch prediction on CPU very difficult and inefficient.
The alternative is to not only jit the instructions but also embed pieces of the interpreter between them so that the compiler can see how those instructions are connected and generate the code for the whole sequence of them. This the compiler can make better assumptions about the code and optimize it a lot more.
This work eventually made it to all sorts of virtual machines. Adobe used it for Flash, Mozilla for their SpiderMonkey javascript engine, Google used this design for early versions of Android Dalvik VM.
PyPy uses it, too. That's why they call it a meta-compiler. It compiles pieces of your code and pieces of its own interpreter together to produce the more optimized binary.
It's not the dispatch to the next instruction that's expensive in most virtual machines. It's the sheer complexity of each instruction as it maps to the underlying assembly instructions [2]
Just getting rid of the dispatch loop doesn't help much and in many cases the increased pressure on the instruction cache makes performance worse.
I'm trying to help with this problem at the moment for the CRuby JIT compiler.
SpiderMonkey, LuaJIT and Dalvik record the trace in the bytecode interpreter. There's no generating baseline JITed code with additional trace recording stuff.
What PyPy does is quite different. PyPy is a Python virtual machine written in a language called RPython and has an interpreter and trace compiler for RPython.
It adds a whole layer of indirection compared to traditional trace compilers which makes it harder to do some optimizations but makes it easier to implement some more basic parts of trace recording and compilation and crucially, makes it somewhat re-usable for different languages.
[1] https://archive.org/details/optimizationofho00fish/mode/2up [2] https://www.sciencedirect.com/science/article/pii/S157106610...
Also it doesn't speed up things that rely on external C libraries, so code using numpy/scipy/tensorflow/etc doesn't generally run appreciably faster.
If we put this figure in context with the CLBG (see https://benchmarksgame-team.pages.debian.net/benchmarksgame/...) we're somewhere between Racket and Dart.
Is a further speedup to be expected? Why is it still slower than V8 after so much development effort?
Anything more realistic involving actual object access, not even allocation, and V8 wins by large margins.
time luajit-2.1.0-beta3 nbody.lua 50000000 real 0m8.437s
time node nbody.js 50000000 real 0m5.065s
I don't know what I wanted to say with that, but I know pypy has spent a lot of time rewriting some core libraries in python/Rpython. I thought it was to make pypy more hackable, but maybe it has some jit benefits as well.