Implementing Fast Interpreters
nominolo.blogspot.com
nominolo.blogspot.com
For an example see: http://lambda-the-ultimate.org/node/3851 (he goes by MikePall - without spaces)
[here is my port of DeltaBlue: https://github.com/mraleph/deltablue.lua, I admit that I might have screwed up porting it, but internal benchmark verification checks did not catch anything]
I'm not sure this makes sense as proposed. Where is the CPU going to put these pre-decoded instructions? The instruction decoder is the first stage of a pipeline, and the decoded instructions are usually fed directly to subsequent pipeline stages. When a branch is predicted, the stream of instructions continues to be decoded and executed in program order from the predicted branch target, and the pipeline stays full. But you can't use this "PREDECODE" instruction to keep the pipeline full, because you can't speculatively execute instructions unless you have already executed all of their dependent instructions (ie. all the instructions prior to the branch that will jump to the target we are feeing to PREDECODE).
What would make sense (to me at least) is if you could have an instruction that tells the branch predictor "the branch at address X will most likely jump to address Y next time." The predictor could then update its tables to adjust the prediction it will make. It seems like this should be pretty straightforward; since the branch predictor and instruction decoder both live at the head of the pipeline, there shouldn't be any danger that the hint is registered only after the branch has already been predicted.
I don't know what folks do now, but Intel had a trace cache as of a few years ago. That trace cache contains decoded instructions.
> What would make sense (to me at least) is if you could have an instruction that tells the branch predictor "the branch at address X will most likely jump to address Y next time."
US Patent 5,949,995 .
http://morepypy.blogspot.com/2011/04/tutorial-writing-interp...
http://morepypy.blogspot.com/2011/04/tutorial-part-2-adding-...
http://tratt.net/laurie/tech_articles/articles/fast_enough_v...
http://tratt.net/laurie/research/talks/2012/kent_fast_enough...
That said, in a language like Python, the interpretation overhead matters much less, because bytecodes are typically very "fat" and involve say multiple dictionary lookups (which are way more costly than bytecode dispatch lookup). If I were to optimize the interpreter, I would start with specialized bytecodes, that avoids fatness of the bytecodes first.
Is that number 4000K correct? That sounds awfully big for me.
4115 buildvm_arm.dasc
4873 buildvm_ppc.dasc
3704 buildvm_ppcspe.dasc
6458 buildvm_x86.dasc
x86 and x86-64 share code.Actually, if you're reading the original linked article, you should also be reading Andy's series on V8.
[1] http://wingolog.org/archives/2012/06/27/inside-javascriptcor...
V8 does not have an interpreter. Lithium is V8's (low level) intermediate representation.