So you could conceivably afford a wide frontend if you restricted x86 to a subset (64 bit only, drop a bunch of weird CISC'y instructions).
So you could conceivably afford a wide frontend if you restricted x86 to a subset (64 bit only, drop a bunch of weird CISC'y instructions).
IIRC Jim Keller said in some interview that modern x86 uses prediction and speculation (similar to branch prediction), and it works surprisingly well.
But fundamentally, a given chip, dedicating a given area to the task, can only begin to decode at so many positions per cycle. And the more intelligent it tries to be about where to start decoding, the longer into that cycle it needs to wait.
And one nastiness about x86 is that you have to decode pretty far into an instruction to even determine its likely length. You can’t do something like looking up the likely length in a table indexed by the first byte of an instruction.
I wonder whether modern chips have pipelined decoders.