It’s mainly (I think) the distribution of instruction lengths. x86 instructions have any length from 1 to 15 bytes. A decoder wants to decode multiple instructions per cycle, and it generally does this in parallel, by simultaneously deciding at multiple starting points. With a fixed length ISA, to decode n instructions, you just decode them. With x86, if you simultaneously decode at offset 0, 1, …, 7, you have 8 decoders but are only likely to decode a couple of correct instructions. The rest start in the middle of an instruction and need to be discarded. So you either need many more parallel decoders for the same throughput or a more complex system to try to avoid throwing away so much work.