How to Read ARM64 Assembly Language
wolchok.org
wolchok.org
I'm curious (I know very little about assembly or what's on the CPU so pardon me if what I'm asking makes no sense), what's the benefit of having a whole instruction that both multiplies and adds? Is there a logical gate on the processor that does this? Or is this just going through a binary multiplier before going through an adder? Does what I'm asking even make any sense?
Multiply-add is a good choice because it corresponds to the relatively common operation of computing the address of a field of a struct in an array so you can operate on that field.
(e.g. &(points[5].x) is &points + (5 * sizeof(point)) + offsetof(point, x)).
There are a fair number of insns in the A64 instruction set that make use of this trick to provide one flexible instruction that as a special case provides useful simpler functionality under an alias. (Register-to-register 'mov' being an alias of 'orr' is another.)
This is only relatively common inside loops. Inside loops you will usually index with the loop counter or some other value that is derived from it linearly. Compilers will typically use induction variable arithmetic that doesn't involve multiplication.
It’s an ALU, way more complex than a logic gate (of which it’s composed), but yes fused multiply-add units are standard on every modern CPU. In fact if your processor is recent (more so than Haswell) odds are good it only has FMA FP ALU, no pure adder or multiplier.
In general, outwith ALUs as well as in, it is very cheap to fold any (reasonable) number of additions and subtractions, even ones with constant left/right shifts/rotates to the addends and subtrahends, into multipliers.
> Scottish Twitter users 'shocked' after discovering the word 'outwith' is only used in Scotland [0]
[0] https://www.dailyrecord.co.uk/scotland-now/scottish-twitter-...
Having looked in the ARM reference manual, the "MUL" instruction is just an alias for MADD with an addition of zero!
I can't find timings for this instruction with 30 seconds of googling, has anyone got a spec with instruction timings?
Source: https://dougallj.github.io/applecpu/firestorm-simd.html
It's still commonly said that RISC processors are faster than CISC because they are "reduced", as in they have fewer instructions. But really it's very beneficial to add instructions that do a lot, if it's something that can easily be done in hardware and replaces several simpler ones.
Multiply-add is an example of one; others are bitfield extraction and rotation, SIMD shuffle, AES encryption, and some of the complex memory operands x86 and ARM have. I even still think x86's memcpy instruction is a good idea.
struct X
{
float a;
int b[10];
};
X x;
If you want to access x.b[3], then you have to add sizeof float to the address of x, and then add sizeof int times 3.Ironically, the best optimization that isn't done to perilogues in the x86 architecture is to take out the PUSH/POP instructions and replace them with MOVs and LEAs, ironically making them more like perilogues on other ISAs.
What? In x86 asm notation, destination is always on the left.
EDIT: I've only ever used the notation defined by the Intel Programmer's Guide official documentation. My bad.
You can run some of it (asm64 work but not the objdump or gdb, only 32 bit?) by using the docker under macOS
``` docker run -it --entrypoint "/bin/bash" ubuntu:latest ```
```
apt install qemu-user qemu-user-static gcc-aarch64-linux-gnu binutils-aarch64-linux-gnu binutils-aarch64-linux-gnu-dbg build-essential
apt install vim
apt install arm-linux-gnueabihf
apt install gdb-multiarch
apt install arm-linux-gnueabihf-gcc
apt install g++-aarch64-linux-gnu
```
For the vector source, I wonder why use c++ for testing instead of just c. After adding cout and int main(), the program can run. However, as the linked one seems not work for arm64 so far, I am still somewhere in between and not able to move to the reading assembly part, as the target of all test.
Cannot get the ld to work but use a simpler c source can generate something like this:
```
cat vector2c.S .arch armv8-a .file "vector2.c" .text .align 2 .global normSquared .type normSquared, %function normSquared: .LFB0: .cfi_startproc sub sp, sp, #16 .cfi_def_cfa_offset 16 stp x0, x1, [sp] ldr x1, [sp] ldr x0, [sp] mul x1, x1, x0 ldr x2, [sp, 8] ldr x0, [sp, 8] mul x0, x2, x0 add x0, x1, x0 add sp, sp, 16 .cfi_def_cfa_offset 0 ret .cfi_endproc .LFE0: .size normSquared, .-normSquared .align 2 .global main .type main, %function main: .LFB1: .cfi_startproc stp x29, x30, [sp, -32]! .cfi_def_cfa_offset 32 .cfi_offset 29, -32 .cfi_offset 30, -24 mov x29, sp mov x0, 42 str x0, [sp, 16] mov x0, 12 str x0, [sp, 24] ldp x0, x1, [sp, 16] bl normSquared mov w0, 0 ldp x29, x30, [sp], 32 .cfi_restore 30 .cfi_restore 29 .cfi_def_cfa_offset 0 ret .cfi_endproc .LFE1: .size main, .-main .ident "GCC: (Ubuntu 9.3.0-17ubuntu1~20.04) 9.3.0" .section .note.GNU-stack,"",@progbits
```
from vector2.c
```
#include <stdint.h>
struct Vec2 { int64_t x; int64_t y; };
int64_t normSquared(struct Vec2 v) { return v.x * v.x + v.y * v.y; }
int main() { struct Vec2 v; v.x = 42; v.y = 12; normSquared(v); return 0;
}
```
``` #include <cstdint>
struct Vec2 { int64_t x; int64_t y; };
int64_t normSquared(Vec2 v) { return v.x * v.x + v.y * v.y; }
int main() { Vec2 v; v.x = 42; v.y = 10;
int64_t x=0; x = normSquared(v); return 0; }
// gcc vector.cpp -o vector -Wa,-adhln=vectorO0.s -g -march=native //gcc -ggdb3 -o vectorgdb vector.cpp //gdb vectorgdb
```
Getting the interaction working and now can be back to reading his post!
The mechanics of branch-with-link can be explained without using x86 as a base. It's a call where the return address is saved in a register and code controls where and when that address is spilled to the stack, rather than it always being on the stack. This is common to several ISAs.
The explanation that sp is a "stack pointer" is like pretty much every stack-based ISA, and does not need special reference to the x86. The idea that all instructions are the same width, similarly, is common to several ISAs, and does not need special reference to only one of the architectures where it is not the case.
And operand order is not unlike x86, but rather unlike a specific assembly language for x86, for which there are alternatives.