I think instruction density has quite some significance here too - x86 opcodes vary between 1 and 15 bytes with 2-3 being average and ARM has Thumb mode where instructions are either 2 or 4 bytes, but all MIPS instructions are 4 bytes. It also has twice as much L1 as most of the ARM and x86 processors, which apparently didn't help it much. Cache consumes power too, and thus I believe small variable-length encodings (like x86) are ultimately better since they allow for better utilisation of cache; the extra complexity in the decoder to handle this, which basically amounts to a few barrel shifters, is almost nothing in comparison to the area and power that more cache would need.
The entire reason CISC architectures emphasized complex multi-cycle instruction execution is because memory accesses were orders of magnitude slower than the processor and data storage was extremely limited.
When considering cache, these points are all true again. There's a common belief about optimising for x86 to avoid the smaller but slower "CISC" instructions, but in situations like tight loops, an instruction that's 2-3x slower individually can be better than the faster longer one(s) if it means the difference between code and data staying in cache or a 10x+ slowdown from a cache miss somewhere else. Especially on an OoO/superscalar design where the slower instruction can be executed in parallel with other nondependent ones. (Intel/AMD's focus on speeding up these small CISC instructions - which they have done - is possibly one of the reasons why x86 performance continues to improve.)