I suppose, for twice the latency, CPU designers could implement (a hash of) the previous branch address from the source instruction pointer as part of the BTB tag, similar to using the global branch history as part of the state in the branch predictor. Presumably, the global branch history could also be used in the BTB tag to give some hint as to which bytecode we've just finished executing.
Though, where it really counts, interpreter writers are already often using computed gotos, reducing the reward to cost ratio for implementing such specialized BTB improvements.
On the other hand,
while (1) {
switch (...) {
...
}
}
is probably rare enough (and almost certainly in a hot loop) that a very specific optimization flag might be better than a syntactic extension. Granted, an optimization flag doesn't work for Erlang's threaded code use case.I use them in both my C befunge implementations:
https://github.com/serprex/Befunge/blob/c97c8e63a4eb262f3a60...
https://github.com/serprex/Befunge/blob/c97c8e63a4eb262f3a60...