https://debugger.medium.com/why-is-apples-m1-chip-so-fast-32...
https://debugger.medium.com/why-is-apples-m1-chip-so-fast-32...
It seems unwise to draw far-reaching conclusions about RISC-V or even ARM64's intrinsic merits versus Apple's CPU designers when there are so many variables. The frontend decoder hasn't been a frequent bottleneck in Intel cores for a long time and they could scale it up more aggressively if they wanted.
Apple's engineers did a great job. That seems to be the conclusion we can draw based on currently available evidence.
This isn't grounded in any facts. Decoding the variable length x86 ISA costs you exponentially in decoding width, both power and area. You can scale it, but it will never be efficient. The way Intel and AMD combat this is by having a decoded uOP cache from which the issue width is typically twice that of the frontend decoder. Arm64 has an inherent advantage here (RISC-V does not have quite the same advantage as RV64GC instructions are a mix of 16- and 32-bit). Arm64 also is much more recent design than x86_64 that learned a lot from the past experience and isn't bogged down by a lot of useless legacy. This helps.
Arm64 is rather large for a RISC ISA, but it's mostly pretty good (however IMO RISC-V's lack of flags and implementation of conditional branches is superior).
All the legacy junk in x86 is obviously a pain for Intel to support. Any blank-slate ISA is going to have an advantage there.
Most classic computationally intensive work (video encoding, science, but also benchmarks) spends its time in fairly tight loops or small kernels, running over large data. uop caches make decode bandwidth irrelevant here.
But general usage of a machine sees the instruction pointer wander all over the place (particularly if you have multiple tabs of JavaScript open). More decode bandwidth means more performance here.
Are compilers are an an example of a heavy workload with a large hot code size? It would be interesting to compare the M1's advantage in compiling to its advantage in, say, video encoding.
But it does mean that x86 designs tend to carefully balance the size of the decoders to other structures to make sure they're not the binding constraint too often. With ARM the approach seems to be more to make the front end 50% bigger than you think you need to be sure it's never a problem and refill the front end buffers more quickly after a mispredict.
[1] The Hillis-Steele paper on data-parallel algorithms from 1986 describes this algorithm for parallel lexing.
It simply isn't practically scalable much beyond where we are; if it were, you can be sure Intel would have scaled it instead of using µOP caches.
I don't think anyone is saying they could scale up the decoder "for free". If they had a fixed-length ISA, I'm sure they would have increased the decoder width sooner (and using different techniques) since with high-end out-of-order cores you're always looking for cheap ways to over-provision your pipeline even if it only helps on some workloads some of the time. Their current use of the uop cache tells us that they consider it the most economical trade-off at that point in the design space (where the decoder can output up to 4 instructions and the uop cache can output up to 6 instructions); you can't infer that they've hit an impassable brick-wall with instruction decoding.
But of course if the target isn't a high end application processor but instead a microcontroller, say, RISC-V's simplicity has a lot going for it. Or for a grad student implementing a simple OoO processor in a semester long class. Or back when I was doing my thesis having an open source core to modify would have been a huge advantage. As the article says RISC-V can be a success without replacing ARM, POWER, and x86.
There's some more discussion in here about the source of the M1's performance, and it largely seems to come down to the smaller process size that enabled Apple to scale up a lot of the structures in the uarch:
I believe this is covered in the medium article that was linked, in part of the discussion on the number of decoders x86 processors have vs the m1. It is in fact this process of breaking instructions into uops that seems to be the bottleneck, and it is apparently not easy to improve this due to the complexities introduced by variable length instructions. If you have reason to think that part of the article is wrong I'm interested in hearing it, I'm not an expert on modern day processor architecture techniques so I don't really have an insider perspective on this issue.
This is also the reason why the whole CISC and RISC debate in it's original form is outdated. The processors internally are all RISC. But the ISA can be more complex.
The x86 ISA makes the decoder harder to parallelize, so it takes more chip area compared to equivalent width for ARM64. And the wider you want to go the harder it becomes, whereas with ARM64 you just slap more decoders.
Another is the x86 memory model that restricts how stores can be issued into memory so that they're visible to other cores.
This is also a good thing for AMD. They could "just" make a Zen ARM CPU. Sure it would be a lot of work, but vast majority of the chip is shared.
The more 1-to-1 your translations to uops are, the less power and die area you need to spend decoding them. In addition, less complex translations means fewer pipeline stages needed for the same design which also has lots of ramifications.
On the other hand if your ISA is the micro-ops directly then the instructions start to take ridiculous amounts of space. It's a balancing act between instruction size (to save instruction cache) and the complexity of decoding them.
And it's not just about being 1:1. It's also about how wide you can go. And variable length encodings simply are fundamentally more hard to parse in parallel fashion. That means a wider unit is harder to achieve, needing more space and power.
I think with x86 Intel and AMD are basically brute forcing this by just pulling instructions off a random position and hoping it’s a correct offset, but it’s very inefficient.
But perhaps Intel/AMD can surprise us with a dynamic allocator that runs in the reorder buffer. Or perhaps they can still push the limit one more time with more transistors. Another option would be to implement a fast-path for small instructions, so in effect they would be moving from CISC to RISC but only for parts of the code that need the extra performance.