In the server side of things there's Neoverse, graviton, etc.. which might have "trickle-down" effects but since it's primarily focused on just having more cores rather than faster cores that seems unlikely. And that also hasn't "proven to be better than x86" either, Epyc is a beast and AMD doesn't seem to be slowing down or hitting any limits in scaling up.
Certain ARM implementations (M1, M2) have proven better than certain implementations of x86 (basically all of them). The ARM implementations in question are not available for general purchase. It's only worth the effort to switch if a competitive processor becomes available. It doesn't look like the MS teamup with qualcomm that resulted in the sq1/sq2 procs are that.
In the server space, that's different given that graviton/cavium have proven competitive.
1. They have been locked into using chips from Qualcomm, which are just slower than Apple's ARM designs. 2. Qualcomm has not implemented the hardware necessary to speed up the x86 translation. The basic problem with running x86 applications on ARM is handling differences in the memory model (x86 uses a strong memory model, ARM uses a weak one). This is slow to handle in software, but IIUC Apple customized their hardware to speed up that operation, which puts them in a much better place for emulating x86.
Certainly the floor for the complexity and size of an ARM processor is lower than for x86 - the ISA is smaller, easier to decode, and suffers from fewer silly legacy baggage items.
However, the reality is that all modern out-of-order microarchitectures are fiendishly complex and that at the "top of the game," implementation details at every level of implementation matter more than the fundamental ISA at play. With x86 there's a certain die size and complexity "tax" in terms of instruction decoding and support that majorly affects small designs. For example, an x86 microcontroller would be a bad idea compared to an ARM one, full stop. But, once you've paid that top-line x86 tax and you're building a huge out-of-order microarchitecture, the differences are minimal and how well you build the rest of the CPU's machinery matters much more than the ISA you started with.
One difference is having constant instruction length, which allows highly parallel decoding and a wider machine in general. This is part of what makes Apple's CPUs faster and cannot be replicated by x86-64, ARM32 or RISC-V with compressed instructions.
RISC has had its moments. Way back, it was better than CISC because the simpler instructions allowed a higher clock rate. Then CISC CPUs turned into RISC CPUs with a CISC-to-RISC translation on the front end. With that in place, there was a whole stretch of time where the only real advantage that RISC had over CISC was that RISC didn't have to have a lump of silicon that translates CISC into RISC instructions.
However, now there's enough space on the die for a CPU to have lots of parallel execution units, and the part of it that became really difficult to scale was the CISC-to-RISC translation unit, because each CISC instruction had an unpredictable length, making working out which instructions you can translate tricky and silicon-consuming.
And so RISC has a significant advantage once more, and that is the ability to vastly-simplify the part of the CPU that feeds instructions into the execution pool, compared to a CISC CPU, because the instructions are fixed-length. This allows this part to translate more instructions per clock cycle than a typical CISC CPU can, and this is what gives the improved performance.
> CISC-to-RISC translation on the front end
I would naively assume that this would be an advantage, since you could easily change the hardware used for any CISC instruction, finding better ways to make it faster. The "work unit" is more abstract, so you could throw the whole problem at dedicated silicon. Or, you could remove dedicated silicon, and just have the CIST spit out a list of RISC instructions.
It seems that, for RISC, you could never throw a more abstract "work unit" at dedicated silicon, without buffering instructions, to see if the intent matches the accelerators. Chip specific compilers would almost be required, to handle the abstraction.
In comparison, a RISC instruction decoder knows that each instruction is the same length, so each instruction can be decoded without depending on the ones before it. This simplifies the decoder so much that it makes it possible to decode four instructions per clock cycle without investing in too much silicon to do it, and while keeping a high clock speed.
1. Whenever you do an x86 vs ARM comparison, there's a number of variables that need to be considered, the node size of the CPU, the power envelope, number of cores, cooling, etc... that make it very difficult to do 1 to 1 comparisons.
2. The main issue x86 wise is that every x86 CPU needs support instructions going back to the 1980s for backwards compatibility, which "wastes" a lot of silicon on functions which are rarely used but need to be in there.
3. Apple has the ability to have their products focus on a few specific devices and just one operating system. This helps design CPUs that they know 100% what they need to handle. You couldn't say that a Qualcomm Snapdragon, just because it's ARM, is better than x86.
All that being said, I'm very much someone who prefers x86 hardware and Windows, but the M1/M2 and Rosetta are very impressive pieces of hardware/software that hopefully kick Microsoft/Linux/Intel/AMD to innovate.
Curious, is this due to something about x86 design (eg: technical benefits over ARM), or are you just referring to the "PC" hardware ecosystem in general (as opposed to Apple/macOS)?
It is more performant because the silicon is made by a company that didn't botch their last process node.
It isn't. ARM's fixed-length instructions helped Apple achieve a wider front-end than contemporary x86 CPUs which helped, but that's about the end of the ISA differences.
Which is why all the other ARM CPUs are slower than x86 CPUs. ARM doesn't really provide an advantage. Apple's stonking huge cash and R&D budget along with Intel struggling right as TSMC is firing on all cylinders is what gives M1/M2 an advantage.
They already have that.