Looking at the results I get, I feel like there's also an inlined switch lookup so that the CPU can get ahead of it.
Basically if I had to the same, I would try to inline the top of the switch down into the break; jump directly instead of jumping to a common location.
So now the operator case ends with looks like
op1_offset:
<invariant=pc> (operator code)
pc2 = pc+1
jmp_off = ops_table+pc2
jmp jmp_off
Because there's no jump between the operator and the pc2 = pc+1, the CPU can compute jmp_off during any time from entering the op1_offset (and this is why you have an HLT opcode, also 32 bit opcodes are much more compact than 64 bit direct threaded CGOTO).I'm not sure if the inlining of the jmp to_top_of_loop can be done by the compiler itself or if it is done by the lower levels of the CPU.
Compilers used to ruin the immediate dispatch of computed goto, to extract the common jump code - I've had to force GCC to leave my CGOTO alone to make this work before[1].