What does matter, IMO:
- assembling a killer team
- 5nm process
- high speed, low latency DRAM
- big-little
What does matter, IMO:
- assembling a killer team
- 5nm process
- high speed, low latency DRAM
- big-little
If these are actually the reasons for the performance difference, and it's difficult to do these on x86 because of the instruction set, it seems to this amateur that ARM64 really does have an advantage over x86.
"A weak memory ordering model, like the one in Apple silicon, gives the processor more flexibility to reorder memory instructions and improve performance, but doesn’t add implicit memory barriers."
[1] https://developer.apple.com/documentation/apple_silicon/addr...
Here's a kernel extension someone built to manipulate this feature: https://github.com/saagarjha/TSOEnabler
> Other contemporary designs such as AMD’s Zen(1 through 3) and Intel’s µarch’s, x86 CPUs today still only feature a 4-wide decoder designs (Intel is 1+4) that is seemingly limited from going wider at this point in time due to the ISA’s inherent variable instruction length nature, making designing decoders that are able to deal with aspect of the architecture more difficult compared to the ARM ISA’s fixed-length instructions.
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
While this is wonderful for ARM in the now-term, we just moved from walled ISAs to a plurality of ISAs, compute just became a bulk commodity in a way that it could not with an x86 duopoly.
Anyone can now take off the shelf RISC-V designs that are currently at > 7.1 coremarks/mhz and get them fabbed on Glofo or TSMC. If you need integrator help, you can use the design services of SiFive.
ARM has been improving much faster than Intel.
Apple has been executing ARM much better than anyone else.
Apple's auxiliary processors and integration have been top notch.
TSMC has been crushing Intel in getting to 5nm.