One of the reasons why M1 is good is pure and simple that it has a pretty enormous transistor budget, not solely because it's ARM.
One of the reasons why M1 is good is pure and simple that it has a pretty enormous transistor budget, not solely because it's ARM.
It's also very hard to achieve more than 4X parallelism (though I think Ice Lake got 6X at some additional cost) in decode, making instruction level parallelism harder. X86's hack to get around this is SMT/hyperthreading to keep the core fed with 2X instruction streams, but that adds a lot more complexity and is a security minefield.
Last but not least: ARM's looser default memory model allows for more read/write reordering and a simpler cache.
ARM has a distinct simplicity and low-overhead advantage over X86/X64.
Furthermore, the high-performance ARM designs, starting with the Cortex-A77, started using the same trick---the 6-wide execution happens only when instructions are being fed from the decoded macro-op cache.
I have vTune installed so I guess I could investigate this if I dig out the right PMCs
0 lsd_uops
1,092,318,746 idq_dsb_uops ( +- 0.49% )
4,045,959,682 idq_mite_uops ( +- 0.06% )
The LSD is disabled in this chip (Skylake) due to errata, but we can see only 1/5th of the uops come from the uops cache. However, the more relevant experiment in terms of power is how many cycles is the cache active instead of the decoders: 0 lsd_cycles_active
378,993,057 idq_dsb_cycles ( +- 0.18% )
1,616,999,501 idq_mite_cycles ( +- 0.07% )
The ratio is similar: the regular decoders are not active only around 1/5th of the time.In comparison, gzipping a 20M file looks a lot better:
0 lsd_cycles_active
2,900,847,992 idq_dsb_cycles ( +- 0.07% )
407,705,985 idq_mite_cycles ( +- 0.33% )Forget Bitcoin mining... how many tons of CO2 are released annually decoding the X86 instruction set?
I’d say ARM has a big advantage for instruction level parallelism with 32 registers.
And it seems to me that ARM has an advantage here. If you want execute 8 instructions in parallel, you gotta actually have 8 independent things that need to get executed. I guess you could have a giant out of order buffer, and include stack locations in your register renaming scheme, but it seems much easier to find parallelism if a bunch of adjacent instructions are explicitly independent. Which is much easier if you have more registers - the compiler can then help the cpu keeping all those instruction units fed.
> include stack locations in your register renaming scheme
Registers aren't related to the stack. "The" stack is just RAM being accessed in a specific cache friendly pattern, with additional optimizations (if you use specific registers) from the hardware in the form of the stack engine. The compiler explicitly loads and stores to and from the registers named by the ISA. Register renaming has absolutely nothing to do with the stack.
When the CPU can tell that a later instruction doesn't depend on the previous value of a register, it's free to rename it. The result is that two independent registers get used even though only one was ever directly referenced. In reality, there are a _huge_ number of registers available on modern processors. Estimates place Skylake, Zen, and Cortex-X1 at 200+, with the M1 at 600+. The ISA just doesn't provide a way to access them directly. (If you want to read about this, the term to look up is reorder buffer.)
Also, there is a giant out of order buffer for stores waiting to be written back to L1. That buffer does indeed have to keep track of cache locations, which directly map to memory addresses, which sometimes happen to refer to stack locations. So in a sense, what you suggested already exists. (If you want to read about this, the term to look up is store buffer.)
> it seems much easier to find parallelism if a bunch of adjacent instructions are explicitly independent
That would indeed make things simpler in some cases. However, many operations such as loading a value into a register (ex mov, [addr]) or zeroing it (ex xor eax, eax) explicitly break the dependency chain by definition. Cases where the CPU fails to properly account for this are documented as false dependencies.
> the compiler can then help the cpu keeping all those instruction units fed
The "compiler handles ordering" thing was tried with Itanium. It seems it didn't go so well.
The CPU is free to simultaneously load two different pieces of data into the "same" register and execute two independent instruction streams on that "single" register thanks to renaming. Speculative execution helps when the CPU can't be completely certain that there isn't a dependency.
For particularly complicated sequences, the compiler spilling due to running out of named registers could indeed pose an issue. However, the CPU is free to elide a store followed by a load if it determines that the address is the same. (If you want to read about this, terms to look up include store-to-load forwarding and load-hit-store.)
I know Itanium didn’t work - but that’s because here the compiler is supposed to do all the reordering work. That’s different from allowing the compiler to explicitly define that instructions are independent by having more registers.
Although apparently Zen 2 changed this and can pull off zero latency. (https://www.agner.org/forum/viewtopic.php?t=41)
Some general background: (https://travisdowns.github.io/blog/2019/06/11/speed-limits.h...)
a = m[i+1] + b
c = m[i+3] + c
e = m[i+7] + d
assume you only have 3 registers, in a RISKy architecture. Every statement becomes something like r1 = *pb // load c
r2 = r0[1] // m[m+1]
r1 = r1 + r2 // a = ...
*pa = r1
Since all registers are used, and all but two instructions are dependent, in the assembly the blocks have to follow one another. There`s also spilling of the b,c,d variables, they have to be read from registers (which could be elided). Assuming no re-order buffer, these instructions runs in three cycles (the first two are independent) - even though the top level instructions are independent.If you want them to run all statements with 4 instructions at a time, you need to have a reorder buffer that covers the whole sequence (12 instructions). (Imagine if b,c,d get modified inside the inner loop and spilled into memory, you have to track memory locations in order to do register renaming.)
Now lets assume you have 6 registers. Now all variables fit in registers and the compiler can easily interleave the code giving a sequence of 3 or 4 independent instructions at a time. If you want to run 4 instructions at the same time, you need no reorder buffer.
This is a kind of specific example, but it shows that if you have more registers (i.e. ARM vs x86), the compiler can more easily interleave instructions, which can help reduce the number of instructions that need to be in the reorder buffer. Or with the same size re-order buffer, its easier to find more independent instructions and keep all the execution units fed. Or, when jumping to some code thats not in pipeline or icache, it allows to sooner run more instructions in parallel, when only a small number of instructions are decoded and in the re-order buffer.
In practice, x86_64 works just fine for HPC number crunching code. Outside of some serious number crunching, when are you going to have more live values than named registers, have instruction streams whose output depends on _all_ of those values (which is why they would be live), and also those streams complete so quickly that you stall on the next set of loads? And you have absolutely no other useful work to do? Honestly I think you're being silly.
Historically, I understand that the 32 bit version of x86 did have scheduling challenges surrounding function calls. The 64 bit version of the ISA expanded the number of named registers and (as far as I understand things) it largely resolved the issue.
Also note that typical hardware can sustain a surprisingly large number of loads per clock. You just need to find something useful to do while you wait for the load to complete. In case you really can't there's also SMT. Really though, the PRF and ROB are only so large.
> If you want to run 4 instructions at the same time, you need no reorder buffer.
You always need a reorder buffer if you want to achieve good performance. Among other issues, the compiler can't predict the latency for each load in advance due to caching behavior depending on the runtime state of the full computer system. I previously mentioned Itanium. It's directly relevant here.
> Imagine if b,c,d get modified inside the inner loop and spilled into memory, you have to track memory locations in order to do register renaming.
No. You can't just rename registers any longer. A store to memory means the memory model for the ISA gets involved. Things become significantly more complicated. The store buffer exists specifically to deal with such issues efficiently on an OoO core. Seriously, go read about it. It's astoundingly complicated for any OoO core regardless of the ISA.
> the compiler can more easily interleave instructions, which can help reduce the number of instructions that need to be in the reorder buffer
Unless I have a serious misunderstanding (I don't design hardware, so I might) everything passes through the reorder buffer. Every instruction is speculative until all previous instructions have retired. (https://news.ycombinator.com/item?id=20165289)
What percent of the die is an ARM instruction decoder?
ARM A32/A64 instruction decoding is dramatically simpler -- all instructions are 32 bits wide and word-aligned, so decoding them in parallel is trivial. T32 ("Thumb") is a bit more complex, but still easier than x86.
What nearly everyone uses is a 16 byte buffer aligned to the program counter being fed into the first stage decode. This first stage, yes has to look at each byte offset as if it could be a new instruction, but doesn't have to do full decode. It only finds instruction length information. From there you feed this length information in and do full decode on the byte offsets that represent actual instruction boundaries. That's how you end up with x86 cores with '4 wide decode' despite needing to initially look at each byte.
Now for the efficiencies. Each length decoder for each byte offset isn't symmetric. Only the length decoder at offset 0 in the buffer has to handle everything, and the other length decoders can simply flag "I can't handle this", and the buffer won't be shifted down past where they were on the next cycle and the byte 0 decoder can fix up any goofiness. Because of this, they can
* be stripped out of instructions that aren't really used much anymore if that helps them
* can be stripped of weird cases like handling crazy usages of prefix bytes
* don't have to handle instructions bigger than their portion of the decode buffer. For instance a length decoder starting at byte 12 can't handle more than a 4 byte instruction anyway, so that can simplify it's logic considerably. That means that the simpler length decoders end up feeding into the higher stack up full decoder selection, so some of the overhead cancels out in a nice way.
On top of that, I think that 5% includes pieces like the microcode ROMs. Modern ARM cores almost certainly have (albeit much smaller) microcode ROMs as well to handle the more complex state transitions.
Once again, totally agreed with your main point, but it's closer than what the general public consensus says.
IMO, you would either go towards bitaligned instructions like the iAPX 432 or the Mill, or 16-bit aligned variable width instructions like the s360 and m68k on the CISC side, and ARM Thumb and RV-C on the RISC side.
That being said, you're definitely thinking about it the right way. Modern Istream bandwidth conscious ISAs absolutely (and perhaps unsurprisingly) look at the problem from a constrained, poor man's huffman encoding perspective similar to how UTF-8 was conceived.
Maybe one could come up with an instruction encoding that encodes some number of instructions per cache line. Every time the cpu jumps to a new instruction (at cache line address + index), the whole cache line needs to be loaded into icache anyway, and could get decoded then -> internally they get represented in microcode anyway.
This is also not a good security property since it means you can hide secret instructions in a program by jumping into the middle of innocuous ones.
> ARM A32/A64 instruction decoding is dramatically simpler -- all instructions are 32 bits wide and word-aligned, so decoding them in parallel is trivial. T32 ("Thumb") is a bit more complex, but still easier than x86.
A64 doesn't have a Thumb equivalent, also, and supporting A32/T32 is optional.
I'm not familiar with how ARM's memory model effects the cache design - Source?
There's a lot of brute force, yes, but it's not the only reason. There are lots of smart design decisions as well.
Plus, most of the last decade software is software that runs on some sort of VM or another (be it JVM, CLR, a Javascript engine or even LLVM).
Soon (in years), x86 will only be needed by professionals that are tied to really old software. And those particular needs will probably be satisfied by decent emulation.
There's also that PC & console gaming markets, which are not small and have not made any movements of any kind towards ARM so far.
For example, the M1 has 128 bit wide memory. This has been standard for decades on the desktop(dual channel), but unheard of in cellphones. The M1 also has similar amounts of cache to the new AMD and Intel chips, but thats several times more than the latest snapdragon. Qualcomm also doesn't just design for the latest node. Most of their volume is on cheaper, less dense nodes.
The M1 is one "node" ahead. Apple forked out the cash to get all their chips on TSMC's 5nm process. This is about 2 years of advancement over the 7nm process AMD pays TSMC for. Intel's latest 10nm node is similarly behind TSMC 5nm.
Semiconductors are tricky. Small performance gains take large increases in power. If you play with overclocking, you'll learn power increases quadratically or even cubically with clocks. The mere "2nm" shrink may seem inconsequential, but for these iso-perfomance comparisons(performance@constant-thermals), it is key.
All this to say, you get what you pay for. Chips can get the same performance on TSMC's 5nm node while using 70% of the power as chips on the 7nm node.[1] Compared to TSMC's 10nm (similar to Intel's popular 14nm still in production), 5nm chips can be expected to use ~45% of the power.
Hopefully that shed some light on the M1's biggest advantage for you.
[1] https://images.anandtech.com/doci/15219/wikichip_tsmc_logic_...
The M1 isn't necessarily a win for Arm in general. Other manufacturers weren't competing before and its yet to be seen if they will.
By going on-package there's almost certainly latency advantages in addition to the much-vaunted bandwidth gains.
That's going to pan out to better perf, and likely better power usage as well.