ARM: Pragmatism, Not Purity
cohost.org
cohost.org
Interesting talk on the history of the architecture, including 64-bit, from Arm's lead architect [1]
[1] https://soundcloud.com/university-of-cambridge/a-history-of-...
Which is exactly what my comment was intended to imply. Nothing else.
These days, RISC-V seems to be quickly gaining terrain in academia.
https://www.jwhitham.org//2016/02/risc-instruction-sets-i-ha...
There were reasons not to use these architectures, even when/if open.
"RISC II proved to be much more successful in silicon and in testing outperformed almost all minicomputers on almost all tasks. For instance, performance ranged from 85% of VAX speed to 256% on a variety of loads. RISC II was also benched against the famous Motorola 68000, then considered to be the best commercial chip implementation, and outperformed it by 140% to 420%."
https://en.wikipedia.org/wiki/Berkeley_RISC
[Not really much mention of MIPS performance relative to CISC minis.]
https://en.wikipedia.org/wiki/Stanford_MIPS https://en.wikipedia.org/wiki/R2000_microprocessor
Elegant.
Also the author goes on to praise x86, which is absolutely fucking bananas if you care about not having to rote memorize register names.
I'm calling a function with four word-sized arguments. In SPARC they go in %o0, %o1, %o2, and %o3. Where do they go in x86?
That said, in practice many people use them interchangeably since A64 is the only instruction set available in AArch64.
That's not pragmatism per se. It's purity, but optimizing for something else: wide front ends and dense code to minimize memory overhead. Those are the right things to optimize for on a modern chip.
Somewhat complicated decodes are fine as long as you don't have to do crazy things to guess the width of instructions like you do on x64.
The complexity this adds to decode is a far cry from the brute-force approach x86 requires.
This negligible impact is weighted against the benefits of higher code density; RISC-V is the most dense ISA, for 32 and 64 bit.
ARM's AARCH64 does in contrast have exceptionally poor code density for a 64bit architecture. This translates in much higher requirements for cache sizes and ROM size, which in turn mean larger area and higher power consumption for a given performance target.
ARM's popularity also stems from its licensing, AFAICT. You can license various ARM cores, from the tiniest to the beefiest, to include in your SOC or MCU, mix and match them on the die, etc. This is not true for any reasonably recent x86 architecture, either from Intel or from AMD.
See a tiny intro e.g. at https://lwn.net/Articles/718888/
Specifically, Thumb and Thumb2 density was excellent, and remained unbeaten until quite recently (2022, RISC-V B+Zc).
But then aarch64 did, to everyone's surprise, ditch variable instruction size, and with it lost its code density advantage.
Thumb2 was remarkable for getting so close to the performance of regular Arm code, but it was always still slower for almost any task even with the reduced instruction cache pressure.
Aarch64 very much loses out; Microcontrollers are resource-constrained, so they won't use it.
High performance implementations (like Apple's M1) need to work around code density by making caches larger, or even having a microop cache; this all has transistor count cost, which imposes area cost, power cost and maximum clock cost.
Off the top of my head, it's not practical to fix RISCV's multi-precision arithmetic issues with instruction fusion. That's not a dominant use-case, especially if post-quantum crypto takes off, but it's definitely a place where RISCV underperforms ARM and x86 by a large factor instruction-for-instruction.
Also there are places where ARM's instructions improve code density, like LDMIA, STMDB etc. These of course can't be fixed by fusion.
I'm sure there are other areas.
RISC-V excels in code density.
64-bit RISC-V is the most dense 64bit ISA by a large margin and has been for a long time.
32-bit RISC-V is the most dense 32bit ISA as of the recent B and Zc work; it used to be Thumb2 was ahead.
All of this is without compromising "purity" or needlessly complicating decode; there already is a 8-wide decode implementation (Ascalon, by Jim Keller's team), matching Apple M1/M2 decode width.
(Thumb2 is the most apt comparison here I think, since the instructions I mentioned are Thumb2 instructions.)
It does not help the largest implementations, that favor having a lot of small ops flying, nor the smallest ones, where fusion is unnecessary complexity.
But it might make sense somewhere in the middle.
Ultimately, it does not harm the ISA to be designed with awareness it exists.
There's lots of O(N^2) structures on a OoO chip to support in-flight operations, so if you can fuse ops (or have an ISA that provides common 'compound' operations in single instructions) that can be a big benefit.
For RISC-V the biggest benefit is probably to fuse a small shift with a load, since that is a really common use case (e.g. array indexing), and adding a small shifter to the memory pipe is very cheap. Alternatively, I think with RVA22 the bitmanip extension is part of the profile and IIRC it contains an add with shift instruction. So maybe compilers targeting RVA22 will instead start to use that one instead of assuming fusing will occur?
B extension helps code density considerably. This is why, if the target has B, compilers will favor the shorter forms.
Citation needed.
>is basically 'instruction fusion will fix everything'.
Wherever I've seen people involved with RISC-V talk about fusion[0], the impression I got is the opposite; it is a minor, completely optional gimmick to them.
Aarch64 has many op-codes that take different parameters to become different effective instructions, often in cases where RISC-V would need an extension for one of their counterparts. I think a RISC-V µarch could do the same internally but that would probably require a larger decoding stage.
The evidence so far points to the opposite.
Existing RISC-V µarch from e.g. Andes and SiFive offer performance that matches or beats the ARM cores they're positioned against, with considerably lower power and significantly smaller area.
And they do already have an answer for most of ARM's own lineup, covering from the smallest cores to the higher performing ones, solely excluding the very newest, highest performance targeting cores, where the gap is already under 2 years.
https://userpages.umbc.edu/~vijay/mashey.on.risc.html
So that would be things like Vax that not only have load-store opps, they also have things like indirect loads and stores in a single operation.
The more complex ones (many mem-mem instructions with fancy addressing modes) are VAX, 68020, NS32000.