Basically computers that start looking and acting more like clusters. Smarter memory and caching, zero-copy/fast-copy-on-mutate IPC as first-class citizens. More system primitives like io_uring to facilitate i/o.
The popularity of modern languages that make concurrency easy means more leveraging of all those cores.
The other factor is Apple is using a "better" process node (TSMC 5nm). I put it in quotes because Intel's 10nm and upcoming nodes may be competitive, but Intel's 14nm is what Apple gets to compete against today, right now.
Intel has been defeated in detail.
Intel's 10nm node is out, I'm typing this on one right now. It's competitive in single-core performance against what we've seen from the M1. Graphics and multi-core it gets beat though...
Or do you mean what Apple used to use? (edit: the following is incorrect) It's true Apple never used an Intel 10nm part.
EDIT: I was wrong! Apple has used an Intel 10nm part. Thanks for the correction!
There are still no 10nm parts for the desktop or the high-end/high-TDP laptops anyway afaik.
That's only for 32-bit ARM; for 64-bit ARM, the instruction size is always constant (there's no Thumb/Thumb2/ThumbEE/etc). It won't surprise me at all if Apple's new ARM processor is 64-bit only (being 64-bit only is allowed for ARM, unlike x86), which means that not only the decoder does not have to worry about the 2-byte instructions from Thumb, but also the decoder does not have to worry about the older 32-bit ARM instructions (including the quirky ones like LDM/STM).
That would also explains why Apple can have wide decoders while their competitors can't: these competitors want to keep compatibility with 32-bit ARM, while Apple doesn't care.
"ROB" and "OOO" might have gotten mixed together here.
OOO = Out-of-order. Refers to the fact that the M1 can decode instructions in parallel.
ROB = Re-Order Buffer. Refers to the stage where the parallel instructions get put back "in-order" and "retired."
[1] https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...