It's intriguingly dense. If you're writing an emulator or a bytecode VM, and you can't just slap a JIT in (because you need to ensure that e.g. data races resolve the same way they did on the original machine), it's very hard to come up with a data structure to represent "decoded" code, that is more performant than just keeping the original code around in memory in its packed-stream-of-opcodes form.
One would think, for example, that it would make sense to do the "instruction decode" pass ahead of time, to end up with an array of pointers-to-instructions plus literal-values... but the resulting representation of the code is usually much larger in memory, and so less of it will fit in cache (and it'll also fight the VM interpreter itself for cache-lines.) You might gain from your instruction impls not having to trampoline back to the interpreter (https://en.wikipedia.org/wiki/Threaded_code, basically), but you'll lose in cache coherence.
Really, the best you can do in such a situation is to translate the stream of opcodes to another stream of opcodes, just ones that you can execute more efficiently (i.e. create your own "microcode" ISA for your VM.)
Either way, "a loop that walks/jumps through an in-memory buffer of variable-length CISC opcodes using a byte-granular program-counter pointer register" seems to be an optimum somehow in design space.