ARM Assembly Is Too High Level: ROR and RRX
xlogicx.net
xlogicx.net
The design of the orginal ARM ISA is very interesting historically, mainly informed by many man years of hand-optimising 6502. As such it was quite idiosyncratic, nice to write by hand, and rather awkward for compilers...
There was always a cost. You had to spend those bits in the instruction encoding that could have been used for more registes, more instructions, more operands, or (the x86 choice) the ability to fit more (smaller) instructions in the instruction cache. The existence of this extra thing in the instruction data path meant you needed an extra cycle or two in the execute pipeline for every instruction, not just the ones with shifts. You also had to have the single-cycle barrel shifter implemented in hardware (this is something that smaller microcontrollers used to skip).
In fact that weird ARM shift field is broadly held to have been a mistake. Note that A64 skips it.
Also very interesting is to see how it was originally implemented (a series of blog posts by Ken Shirriff and Dave Mugridge): https://www.righto.com/search/label/arm and https://daveshacks.blogspot.com/2015/12/inside-alu-of-armv1-...
This is like saying SSE is "too high level" because you could just do all the operations independently with scalar math.
[1] Were, anyway. They're surely microcoded on modern processors.
edit: Also, I'm probably unfamiliar with the type of blog post you're satirizing because I don't read many programming blogs these days. I'd rather just read the documentation (and source code), and any other questions I have are best answered by just asking the computer.
Good satire, like all worthwhile endeavors, is hard. Don't let people like me who don't immediately get it stop you from practicing, though :)
0 isn't permitted with ROR.
http://infocenter.arm.com/help/index.jsp?topic=/com.arm.doc....
xlogicx should read the manual one more time.
I have to be a bit more careful though with this one lately, when I'm working on industrial robots vs. the embedded boards I'm used to!
Thanks. I much prefer this format and can't really comprehend why anyone would prefer a video to a well laid out and illustrated blog post.
However your font size is pretty uncomfortable to read for me.
It's alright to have your own preferences, but so do other people. Some people (like me) prefer audiobooks instead of paper books, some people are militantly the other way around. It is what it is :-) Kudos, certainly, to people who share their work in multiple media.
A technical article such as this one is ideally disseminated using a regular web page. Pictures, text, code all in easy view and scrolled conveniently at will.
Compare that to a video which in essence is an auto-scrolling page that can only be paused or slightly slowed down. You wind up pausing, rewinding and skipping parts. Annoying.
If the author doesn't like this they shouldn't look into MIPS because it goes well beyond that. You see, MIPS has a special "R0" register that's always 0 (AARCH 64 does as well by the way) so you can always use it as a placeholder in other instructions.
As such, there's no real MOVE instruction, it's just an assembler mnemonic that assembles down to `OR $target, $src, $R0`. NOP? It's by convention `SLL $R0, $R0, 0` (which has the nice property of being an instruction encoding as "0x00000000"). You want to negate a number? `SUB $target, $R0, $src".
Since all instructions are 32bit wide you can't load a 32bit immediate value in a single instruction, instead the assembler's "LI" mnemonic generates a pair of instructions (LUI/ORI) for large immediate values (ARM prefers PC-relative loads).
You have a whole bunch of mnemonics in MIPS that are just aliases around other instructions. I always thought it was pretty clever.
In summary you can have this assembler listing:
1:
sll $0, $0, 0
or $t0, $t1, $0
li $t0, 0xabcdef
sub $t0, $0, $t1
j 1b
That will disassemble to: nop
move t0,t1
lui t0,0xab
ori t0,t0,0xcdef
j 0x0
neg t0,t1
The only operation here that I would qualify as "high level" is
the reordering of the "neg" instruction into the delay slot (note
that j is no longer the last instruction). Everything else is
very straightforward substitution and if the assembler didn't
support these mnemonics we could implement them with very trivial
macros.Note that even x86 assemblers do that to some extent, for instance "nop" assembles down to an instruction with no side effect (typically `xchg eax, eax`). Furthermore there are a bunch of mnemonics for the same encoding, for instance JAE (jump if above or equal), JNB (jump if not below) and JNC (jump if not carry). Overall instruction encoding is also massively more complicated in x86 (and even more for amd64) so the assembler needs to handle many more corner cases than the simple substitutions of ARM and MIPS. As a brain teaser, consider the following similar looking amd64 instructions that load the 32bit value pointed at by a register (the only difference is that the first one dereferences the pointer in %rax, the second in %r12):
mov %eax, (%rax) ; assembles to 89 00
mov %eax, (%r12) ; assembles to 41 89 04 24
I can't even be bothered to walk you through this but basically
it has to do with the fact that %r12 happens to be encoded as
%rsp + 8 (because registers r8 to r15 are effectively a hack
since x86 only supported 8 GPRs) and %rsp has special semantics
in this addressing mode which mandate a different, longer
encoding otherwise you end up with an ambiguous instruction.Yeah, I think in retrospect we can give ARM a pass for their ROR shenanigans.
Sounds like CISC.
[1] https://svkt.org/~simias/up/20180926-112325_x86-mov.png
[2] https://svkt.org/~simias/up/20180926-112224_or-encoding.png
CISC has different encodings for MOV and OR (and often subtle differences, e.g. OR updates the flags and MOV doesn't). On RISC processors, which have MOV as a special case of OR, the assembler accepts MOV at the source code level but the processor does not have to implement a separate MOV instruction at the binary level. Therefore the instruction set is indeed reduced.
ARM's peculiarity is that the fundamental ALU operation is "R1 op (R2 shiftop R3)" or "R1 op (R2 shiftop #nn". But it's still not a CISC design, it's just that the barrel shifter is at a different place in the ALU and that shows in the instruction encoding. Apart from this quirk the ideas from the previous paragraph apply just as well to ARM.
I still find this classic to be the best explanation of the technical characteristics of CISC/RISC: https://userpages.umbc.edu/~vijay/mashey.on.risc.html
Here's a short summary:
Most RISCs:
- Have 1 size of instruction in an instruction stream
- And that size is 4 bytes
- Have a handful (1-4) addressing modes
- Have NO indirect addressing in any form (i.e., where you need one memory access to get the address of another operand in memory)
- Have NO operations that combine load/store with arithmetic, i.e., like add from memory, or add to memory.
- Have no more than 1 memory-addressed operand per instruction
- Do NOT support arbitrary alignment of data for loads/stores
- Use an MMU for a data address no more than once per instruction
- Have >= 5 bits per integer register specifier
- Have >= 4 bits per FP register specifier
>Have 1 size of instruction in an instruction stream and that size if 4 bytes
So that means that Thumb isn't RISC because it has 16 bits instructions and a few double-width opcodes? Even though its instruction set if effectively even more restricted than ARM? That doesn't make sense to me.
>Do NOT support arbitrary alignment of data for loads/stores
MIPS has SWL/SWR LWL/LWR, does that count? I suppose you could say that RISC has no support for arbitrary alignment in regular load and store instructions but again, is that really enough to disqualify an IA? What if I made a tweaked MIPS CPU with an identical instruction set with the only difference being that unaligned LW/SW would work as intended instead of raising an exception, would it stop being RISC?
>Have >= 5 bits per integer register specifier, Have >= 4 bits per FP register specifier
That actually disqualifies ARM32 as far as I can tell, since it only has 16GPRs encoded using 4 bits. I fail to see how this small encoding detail is relevant to RISC anyway. Maybe it just meas that you need at least 32GPRs?
Wikipedia has a much broader (and IMO more reasonable) definition of RISC:
>Various suggestions have been made regarding a precise definition of RISC, but the general concept is that such a computer has a small set of simple and general instructions, rather than a large set of complex and specialized instructions.
By this definition an instruction such as "Floating-point Javascript Convert to Signed fixed-point, rounding toward Zero" is very much un-risc-y.
Hence, assembling a single mnemonic into two or more instructions is fairly common on RISC architectures. But the instruction set at the machine level is still simple and direct, even if the assembler expands.
As a fun fact, modern x86 will often just microcode complex instructions inside the CPU to several simple mu-ops and then execute those. In turn, they are just doing the same work as the assembler, but in hardware. It is necessary for backwards compatibility, but it is hardly elegant.
ror r0, #0
gets assembled to
mov r0, r0.
I see how both are nops and assume they affect state congruently, but why is one nop encoding preferred over another? Does the architecture do something desirable when encountering the MOV incarnation vs. any other?
The ARM Instruction Set PDF I'm looking at only lists MOV as a real instruction -- one with a distinct opcode -- out of the above three.
As a note, I'm basing this off of the v7-A and v7-R manual.
That manual uses unified assembly syntax, meaning that old ARM and Thumb are described together. Because of irregularities in Thumb, the old ARM instructions end up being described in a way that is needlessly verbose. You lose the insight into how the opcodes are actually decoded.
Look at an older manual. ARMv5 will do nicely. There, you can see that the MOV instruction is mostly described by 2 bits. (one condition code is stolen) In an even older ARM, such as ARMv4 I think, it really is just 2 bits.
On a higher level, this post got me to realize that there could be ops (in this case nops) that have equivalent effects but are encoded as different instructions. Even though registers might state-change equivalently, I can see that the internal processor state might mutate differently. Maybe one nop encoding is faster, and maybe the caches get hit differently, etc. As an assembler, what reasons might there be to prefer one encoding over another?
When designing an architecture, I can imagine putting a "fast nop" in the instruction encoder that essentially just short circuits around it. Is this something that's done in practice?
For example:
ldr rn,=0xCAFEBABE
Will actually be assembled as: ldr rn,[pc,#(literals - .)]
literals:
.word 0xCAFEBABE
IMO it really is odd that some instruction to the assembler are not direct translations. In theory it can cause some issues if the developer makes irresponsible assumptions based on the program text. The mnemonics make life easier for folks who are familiar with the assembler, but harder for someone poking around doing some detective work.I tried this IBM one but it has a terrible clickthrough to attempt to avoid GDPR:
https://www.ibm.com/developerworks/library/l-ppc/index.html
And, as seems to pretty much always be the case, the wikibooks looks promising but then appears to be empty:
https://en.wikibooks.org/wiki/PowerPC_Assembly/Instructions
This one seemed pretty good:
https://www.cs.uaf.edu/2011/fall/cs301/lecture/11_21_PowerPC...
https://blogs.msdn.microsoft.com/oldnewthing/20180806-00/?p=...