The Fastest VM Bytecode Interpreter (2010)
byteworm.com
byteworm.com
for (pr = pn, ip = i;; ip++) {
op = &optable[ip->op];
...That looks like a switch-loop interpreter, which is considerably slower than a direct threaded interpreter.
GCC supports goto * * which is a special mechanism that allows for going even faster for the inner loop.
You can see a very similar comparison, comparing the Mono JIT vs an interpreter from the early 2000s here - http://rhysweatherley.sys-con.com/node/38831
You should also try to play nice with the branch prediction machinery. Doing unbalanced rets isn't good. Doing all the jumps or calls from a single jump or call site isn't good, either. Spread them around so you only have one destination per call site. The method in the above paper is pretty much the ultimate you can easily do without rolling out the really heavy machinery.
I agree that the branch (target) prediction will fail horribly on a naive interpreter that calls different branch targets from the same static instruction over and over again. That alone explains the low performance.
Also, am I wrong that, at least in most Unixish systems, interrupts use a separate stack register from that of user space code, and thus, with care, one can avoid corruption of the direct thread addresses?
http://en.wikipedia.org/wiki/Branch_predictor#Prediction_of_...
http://infocenter.arm.com/help/index.jsp?topic=/com.arm.doc....
This type of branch predictor went mainstream in the late 90's with CPUs such as the AMD K6-2.
It is not a correctness issue (as long as there's nothing else that thinks it can put stuff on the stack), just a performance issue.
You should not expect that the code snippets you "return" to will leave the stack alone -- spill/fill and callee-save instructions etc will kill the part of the "stack" that has already been "returned" from which is fine as long as you don't want to reuse it. But I think you will.
https://github.com/mono/mono/blob/e85609a84907dc3a919bac289a...
1) His friend did NOT write a bytecode interpretor, he wrote a bytecode compiler: there is a HUGE difference between those two.
2) No, it wasn't faster than assembly code. And the arrogance and audacity to make such a claim is obnoxious. Unless you gained access to Intel microcode, this is just plain wrong.
I wish I could downvote this: its just bad for aspiring compiler/interpreter developers. If you want some quality discussion see Mike Pall: http://article.gmane.org/gmane.comp.lang.lua.general/75426