1) Use a register-based VM (with a sliding and growing register file) instead of a stack-based VM. In theory you can make a stack-based VM fast with lots of macroinstructions that fuse smaller operations together, but it isn't worth it.
2) Use inline caching for method calls, property accesses, and primitive operations that do type checks. In an interpreter you can modify the instruction stream even on platforms that disallow modification of executable code. I know this isn't the origin of the technique in bytecode interpreters, but here's a paper describing it in case it's not obvious:
http://www.lirmm.fr/~ducour/Doc-objets/ECOOP10/papers/6183/6...
3) Pick your value encoding carefully. You almost always want fast immediate integers. On 64-bit platforms it is quite common these days to repurpose some of the NaN range in IEEE doubles for type tags to enable storing doubles in immediate values.
4) Write your interpreter in assembly. Compilers generate terrible code for interpreters, even (especially?) with the use of computed goto / labels-as-values extensions. The register allocators of traditional compilers are designed to optimize loops by moving spill code outside of them and to reduce the impact of function calls. They will not be able to realistically allocate registers across different instruction bodies, and they won't be able to make the correct tradeoff about how much work to push into the slow path of instruction bodies.
5) Rearrange your instruction bodies based on execution / transition frequencies to improve instruction cache performance.
6) Pay close attention to the boundaries between your interpreter and the runtime libraries / the FFI. You don't want to take a bigger hit than you need to every time you call out to native code.