Also note frequency is only part of the whole picture. CAS latency is important as well, which is much higher with Apple.
Also note frequency is only part of the whole picture. CAS latency is important as well, which is much higher with Apple.
>On the cache hierarchy side of things, we’ve known for a long time that Apple’s designs are monstrous, and the A14 Firestorm cores continue this trend. Last year we had speculated that the A13 had 128KB L1 Instruction cache, similar to the 128KB L1 Data cache for which we can test for, however following Darwin kernel source dumps Apple has confirmed that it’s actually a massive 192KB instruction cache.
That’s absolutely enormous and is 3x larger than the competing Arm designs, and 6x larger than current x86 designs, which yet again might explain why Apple does extremely well in very high instruction pressure workloads, such as the popular JavaScript benchmarks.
The huge caches also appear to be extremely fast – the L1D lands in at a 3-cycle load-use latency. AMD has a 32KB 4-cycle cache, whilst Intel’s latest Sunny Cove saw a regression to 5 cycles when they grew the size to 48KB.
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
x86 needs to find a way to scale decoders without blowing the power budget. Given that the decoders are already bigger than the integer units, I suspect that will be a hard thing to do.
For whatever reason, the overall memory system on the M1 systems just seems better than intel. I really wish I could follow more details from on-die cache to how memory is actually loaded / unloaded to speeds, but every time I've looked at it a little it just seems the M1 / Apple are doing it better across the whole stack.