“csinc”, the AArch64 instruction you didn’t know you wanted
danlark.org
danlark.org
You can read about it in https://community.arm.com/arm-community-blogs/b/infrastructu...
https://www.oreilly.com/library/view/hackers-delight-second/...
It should significantly improve the performance on ARM.
I remember talking to an ARM engineer easily 10 years ago and he told us in that nice british accent: "You know, NEON is like 'back in the yard'" :-D. This has changed a lot, but not enough from what you wrote... Bit sad that these SIMD optimizations are still hand written...
In my experience using a 512 wide movemask (to uint64_t) is the fastest on both x86 and arm64. (Edit: just yo clarify, I meant the fastest for iteration, things like SwissMap are better off using 128 wide movemask)
With rvv you don't really what to go from a vector mask to a general purpose non vector register, because the vector length may vary. But I found it really useful that vector masks are always packed into v0. So even with LMUL=8, you can just to a vmseq, switch to LMUL=1 and use vfirst & vmsif & vmandn to iterate through all indices. (Alternatively vfirst & vmsof & vmclr would also work, I'm not sure which one would be faster)
As I understand it they didn't carry predication across from ARM32 to ARM64 for various performance reasons (if you want to be able to re-order instructions, or even agressively pipeline them, you don't want them depending on the result of the immediately-prior instructon).
Predication everywhere (i.e. orthogonal to the rest of the instruction set, and not special-cased) is certainly more RISC than CISC - but having removed it in general, bringing it back for a few specific instructions is arguably CISCy.
Yes, the IT (if-then) instruction (prefix). It is not supported by Cortex-M0, Cortex-M0+, and Cortex-M1, though. Those are the smallest T32 (Thumb-2) microcontroller designs ARM has.
IT can be followed by up to 4 instructions and encodes their predicate bits would have been in 32-bit ARM code (A32). There is not total freedom regarding their predicate bits: they all have to share the same ground condition (3 bits) and then get an individual bit that says whether to execute when that condition is met or when it isn't. The IT instruction is a 16-bit instruction that devotes 8 bits to this -- not 7, because the encoding is weird.
(They don't have branch predictors, either.)
ARM's 32 bit ISA was very regular and it mostly had a single memory access per instruction, but there were some like store multiple which could potentially save every register to memory. By getting rid of that in A64 and replacing it with an instruction to concatenated two registers and stored them in a single memory access they ended up in a far RISCier place than A32.
I got curious about how RISC-V handles this, but only curious enough to find [1] and not dig any further. That answer is from a year ago, so perhaps there have been changes.
There is a new proposal: Zicond, but it is quite crude, with two instructions. The "czero.eqz" instruction does:
rd = (rs2 == 0) ? 0 : rs1;
And the other "czero.nez" tests for "rs2 != 0". Both are supposed to be result in an operand for another instruction, where a zero operand makes it a nop: for conditional add,sub,xor, etc.
Conditional move, however, takes three instructions: two results where either is zero which get or'ed together.https://github.com/riscv/riscv-zicond/blob/main/zicondops.ad...
Otherwise, the intention was that bigger RISC-V cores would detect a conditional branch over a single instruction in the decoder and perform macro-op fusion into a conditional instruction.
This seems like an overhead compared to actually having the instruction available. Could anyone say how material an overhead this is?
> […] It is because of a special feature of the U74 that when it sees a short forward branch over exactly one ALU instruction it pairs the two instructions together in the A and B pipelines and instead of predicting whether the branch in the A pipe is taken or not it uses the result of the comparison to predicate the instruction in the B pipe.
> It turns it into a NOP at the last moment, or doesn't write the result back to the destination register or something like that.
Also note that the compressed relative branch instructions only use 16 bytes to encode.
[0] https://www.reddit.com/r/RISCV/comments/132s19s/hand_optimis...
Bits
I’ve worked a lot on the kernel and I’m no stranger to optimized code. This is still really weirdly written, and in fact the assembly is much more readable, which is funny.
I know clang needs a lot of prodding to output good code (compared to gcc), but I’m curious whether even clang really needs the logic to be so warped.
if (is_pow2_or_zero(len)) {
int grown = len ? len*2 : 1;
ptr = realloc(ptr, (size_t)grown * sizeof *ptr);
}
compiling into this sort of disassembly to calculate the value of grown: lsl w8, w19, #1 // w8 = len*2
cmp w19, #0x0 // is len zero?
csinc w8, w8, wzr, ne // w8 = (w8 if len != 0) or (0+1 if len == 0)
Pretty clever to create that 1 constant using csinc on the wzr zero register. cmp wzr, w19 // set the carry flag if w19 is zero
adc w8, w19, w19 // w8 = w19 + w19 + carry add eax, eax
setz al add eax, eax
jnz skipinc
inc eax
skipinc:
Also 5 bytes.Sorry if this comment is overly pedantic, I just enjoy having an excuse to talk about assembly.
It's worth noting that 0x80000000 would pass this "is zero" check. (I think this is probably a legal compiler optimisation because signed integer overflow is undefined, but I'm not 100% sure either way.)
Using a jump is also a bit risky - slightly better if it's predictable, much worse if it's unpredictable.
As far as size, this is 5 bytes on 32-bit x86 (as stated), 6 bytes on 64-bit x86, but can be 8 bytes if different registers are used:
4501C0 add r8d,r8d
7503 jnz 0x8
41FFC0 inc r8d
(And, unlike the ARM code, you'd need an additional mov instruction if you wanted to preserve the input value.)It feels like an ADC-based variant might be possible on x86 too - CMP and ADC are also x86 instructions. The problem is that ARM and x86 invert the value of the carry flag on subtraction (and comparison), so it doesn't translate directly, and I can't immediately see how to fix it up without using more instructions.
And of course ARM 32 had conditional execution for all instructions. These appear the variants that were useful enough to keep around when the general feature was removed from aarch64
This is the inner loop of multiplication and was very nice to use, but died in the AArch64 transition.
If you are testing quite predictable things, you almost always want to use branch prediction and not predicated/conditional instructions.
If something is totally unpredictable, let's say a binary search that is looking up random elements in a well balanced heap or tree. Each comparison is very unpredictable. A conditional select would work best there:
item = (val < item->val ? item->left : item->right);
if (val == item->val) ...
You could do your tree walks entirely without branch misses if that first line was a select... But it turns out that is not true. Or it's not necessarily true, depending a few (not uncommon) factors, it can be worse to use a select there.If I download, say, debian-11.7.0-amd64-netinst.iso - does it somehow dynamically adapt to all the different AMD and Intel CPUs and uses the instructions available on the users machine?
In short, distributed binaries tend to use "least common denominator" instructions.
I believe one of the pros touted of Gentoo, where everything is compiled locally, is that all the software uses the CPU to it's fullest potential.
So all new CPU instructions of the last 40 years are pretty much used by nobody?
Of course, I don't think many people are using i386 builds anymore - most people would have switched to the more modern x86_64 long ago.
1. https://www.debian.org/releases/jessie/i386/ch02s01.html.en
2. https://www.debian.org/releases/stretch/i386/ch02s01.html.en
There are new instructions that can be used to write a constant-time SHA-256 hash function. A program that contains megabytes of code, more than a million instructions, may reasonably contain only a few of those instructions, because it contains only one or two small hash functions. The instructions are important and effective, but occur only a handful of times among a million instructions.
If you're running software that's heavily math intensive, or ever gets benchmarked, it's a fairly good bet that they will have either conditionally-installed or conditional-executed code that targets more modern processors.
Debian dropped support with sarge in 2005
Not all, but surely a few.
see https://github.com/bminor/glibc/blob/glibc-2.31/sysdeps/x86_...
You'll find lots of discussions happening around this topic, for example: https://www.phoronix.com/news/Arch-Linux-x86-64-v3-Port-RFC
.Net ahead-of-time compilation (that is, compiling the .net / clr VM byte code into something your CPU can run directly) could (but apparently doesn't?) do CPU-specific optimizations. The JIT compiler, however does do some CPU-specific optimizations.
I think that is the reason GCC will not use it, although it may if you set the target CPU with -mcpu=
Given that AArch64 has/had no 16-bit instruction support, it probably made sense to provide a generalization of a setcond instruction to make use of the encoding space of 32-bit instructions, and that's one of the most obvious (the other ones being cond ? imm : 0 or cond ? imm : reg).
10/10 on the website. Clean simple design and doesn't download 4,124 javascript libraries for the purpose of displaying static content.
Anyway, here's John Mashey, who helped design the MIPS, on RISC v CISC:
Looking forward to your future posts!
The ARM has lots of instructions, each fairly simple. Compare this to an architecture where a single instruction can ① compute the address of its operands in main memory, ② read them, ③ carry out its main operation and eventually ④ write the result to main memory, with most of those steps optional and depending on the arguments supplied.
Maybe it is because the mental model of higher level programmer - for an application programmer, anything that involves writing and reading main memory directly is considered simple, whereas combining a conditional, an increment and a write in a single op sounds "not simple".
Instructions like CMOVBL involve only a small number of CPU registers, nothing else, and can't interact with instructions far ahead or behind them in the instruction stream, or with other cores/threads at all. Very little state. They're simple to reason about, both for the compiler authors, the CPU and the poor developer who's chasing a threading bug.
The world is very different now. RISC won, and it won hard. x86 got a lot more registers, and most of the baroque instructions are no longer used because they got relegated to microcode. As it's used, it's much closer to Berkeley RISC than it is to VAX or S/370. ARM is even more so: the only instructions that touch memory are loads and stores.
As a rough approximation, "simple" instructions can be implemented in a reasonable amount of silicon without microcode, and complex instructions can't.
If you're looking for truly complex instructions, you should look for things like VMENTER or IRET.
(1) various HPC folks have proposed and used extensions that do _part_ of a complex math operation in SIMD at various times, and on GPUs this sort of thing is very common.
(2) except for math libraries that haven't been updated for a couple decades.
(HN seems to have a limit on how deeply you can nest replies. I can't reply to @klelatti's post direct :-/ )
The best approach is always just to verify the disassembly.
Maybe it will come with more time on ARM; but three years in, it's still not there.
There isn't supposed to be some absolute number of instructions that a RISC has, the idea behind it is that you would take a quantitative approach to add instructions, and require that they show benefit beyond a composition of other instructions, and could be practically used.
More capable compilers, wider use of vectorization and other techniques, and more transistors has pushed that a long way since the 1980s.
Complicated things migth be something like division, which can take multiple cycles during which that functional block is busy, or a floating point addition where there are all sorts of complicated rules involving implicit global state around rounding and subnormal and NaNs. Or, most terrifying of all, a load which might target a memory that's been paged out and so require you to bring all the scores of instructions in flight to a halt, switch over to OS code, page in the memory, and then resume as if that one load instruction was the only thing executing at the time the exception happened.
Another example of this are all the combinations with the hard-coded zero register. For instance, the `cmp` "instruction" in A64 (and many other RISC ISAs) is actually an alias to the `subs` (subtract and set status flags) instruction with the zero-register as destination. The idea of the zero register was so potent that modern CISC x86 processors actually have a physical zero register internally, which olde x86 instructions are translated into using.