The M1 can process twice as many instructions per clock cycle than an x86 processor can.
The M1 can process twice as many instructions per clock cycle than an x86 processor can.
Case in point is the Samsung Exynos M3, which despite its 6-wide decoder is barely competitive with the 3-wide ARM Cortex-A76. The Exynos would fare very poorly against any recent 4-wide Intel or AMD CPU.
The crazy thing about the M1 is that every structure in the CPU is huge - the decoders, the ROB, the number of execution units, even the caches. And apparently all of it is implemented very efficiently.
Experience: Designed SPARC, ARM, and x86 CPUs
>The one explanation and theory I have is that Apple might have finally pulled back on their excessive peak power draw at the maximum performance states of the CPUs and GPUs, and thus peak performance wouldn’t have seen such a large jump this generation, but favour more sustainable thermal figures.
Apple’s A12 and A13 chips were large performance upgrades both on the side of the CPU and GPU, however one criticism I had made of the company’s designs is that they both increased the power draw beyond what was usually sustainable in a mobile thermal envelope. This meant that while the designs had amazing peak performance figures, the chips were unable to sustain them for prolonged periods beyond 2-3 minutes. Keeping that in mind, the devices throttled to performance levels that were still ahead of the competition, leaving Apple in a leadership position in terms of efficiency.
https://www.anandtech.com/show/16088/apple-announces-5nm-a14...
It's not as incredibly incorrect as you may think.
An x86 instruction can be as big as 15 bytes and there's no easy way for the decoder to know where one instruction ends and the next one begins.
All ARM instructions are one size, making instruction decoding more efficient and makes out of order processing faster as well.
More details at https://debugger.medium.com/why-is-apples-m1-chip-so-fast-32...
I suspect you haven't internalized this, because reiterating the complexity of decoding instructions just isn't a valid response. The entire point is avoiding that cost.
It could be more accurate to say the M1 can process twice as many half as complex instructions as x86.