The first part of the paper compares token threading vs naive switch in cpython 3.2 which has a compile flag to disable token threading[1]. I were referring to direct threading but the difference is probably minor or irrelevant.
on x64, token threading is
; goto tokens[program[ip]]
mov r0, [program_instructions + ip]
jmp [tokens + r0]
while direct threading is
; goto program[ip]
jmp [program_instructions + ip]
the paper says
Nehalem shows a few outstanding speedups (in the 30 %–
40 % range), as well as Sandy Bridge to a lesser extent,
but the average speedups (geomean of individual speedups)
for Nehalem, Sandy Bridge, and Haswell are respectively
10.1 %, 4.2 %, and 2.8 % with a few outstanding values for
each microarchitecture. The benefits of threaded code decreases
with each new generation of microarchitecture.
but looking at the graphic, it seems a few benchmarks still had a +5% speed boost on Haswell.
It would also have been interesting if the paper provided the generated code by gcc for the switch statement.
If the switch statement's cases are linear (0, 1, 2, 3, ...) and if there are more than 4 cases, then gcc
token threads the code.[2]
[1] https://hg.python.org/cpython/file/v3.3.2/Python/ceval.c#l82...
[2] https://godbolt.org/g/yRjFto