this sgemm progress is pretty impressive. i assume 'version 0' is one of the lower lines on the first graph? he has a graph lower down that shows it getting under 300 megaflops, closer to 50 megaflops for most problem sizes. getting up to 2.2 gigaflops is seriously impressive; that's 55% of the maximum theoretical possible throughput, which is really challenging on things that aren't vector supercomputers (which don't have caches!)
with respect to the rvv assembly, it's worth noting that the allwinner d1 being optimized for here has the pre-ratification version 0.7.1 of the risc-v v extension (rvv) which has some incompatibilities with the ratified 1.0.0. i don't know the instruction set well enough to know if those incompatibilities affect this code
probably worth archiving this link: https://github.com/Zhao-Dongyu/sgemm_riscv
— ⁂ —
vaguely relatedly, i was astounded yesterday to find that gcc 12.2.0 fails at some basic optimizations on risc-v. (all of the following is with -Os)
this trivial subroutine
int sumarray(int *a, int n)
{
int total = 0;
for (int i = 0; i != n; i++) total += a[i];
return total;
}
compiles (with -Os) to a loop containing separate left shift, address-addition, counter-incrementation, and integer addition instructions, 8 instructions in all, plus a ret. omitting initialization: 1: sext.w a3,a5 ; it's my fault i declared i as int instead of size_t so really 7
bne a1,a3,2f ; return if i == n
ret
2: slli a3,a5,2 ; convert i to byte offset
add a3,a3,a4 ; compute &a[i]
lw a3,0(a3) ; load a[i]
addi a5,a5,1 ; increment i
addw a0,a0,a3 ; update total
j 1b ; lather, rinse, repeat
(same in 13.2.0: https://godbolt.org/z/cd1qa6Yzn)i can get it to compile to reasonable code (5 instructions in the loop plus a ret) by writing
int sumarray_opt(int *a, int n)
{
int total = 0;
for (int *end = a + n; a != end; a++) total += *a;
return total;
}
like the desperate, cornered animal i am 1: bne a5,a1,2f
ret
2: lw a4,0(a5)
addi a5,a5,4 ; how motherfucking annoying is it that objdump calls this an add
addw a0,a0,a4
j 1b
isn't strength-reduction something compilers know how to do already? are modern cpus so goddamned superscalar that the difference doesn't matter? (i'm pretty sure less than 0.1% of the billion-plus shipped hardware risc-v cpus out in the field are superscalar.)needing three separate instructions in the first version for converting to a byte offset, adding the byte index to the array base register, and loading from the address is something that could be ameliorated by hardware, but would involve some tradeoffs; the zba bit-manipulation extension to the risc-v isa contains 8 instructions for that sort of thing, including a sh2add that would reduce it to 2. but reducing it to 0 inside the loop with strength reduction is better.
risc-v has a lot of talk about the great importance of macro-op fusion, but that won't help here; it could fuse the shift and add maybe but the add isn't going to go away until somebody applies some loop induction
like, optimizing code for an instruction set with literally thousands of different implementations with different performance characteristics is sort of impossible, but i dare you to find me a risc-v where the first loop runs faster than the second one. or even the same speed
on arm(32) gcc compiles both subroutines to the same code, as any reasonable human being would expect, with the loop being something equivalent to:
1: cmp r3, r1
bxeq lr
ldr r2, [r3], #4
add r0, r0, r2
b 1b
(same in 13.2.0: https://godbolt.org/z/qTs5vMEjs)that is, the same version of gcc does do the strength-reduction optimization for arm(32) (but not aarch64, but maybe that's defensible since aarch64 is usually superscalar: https://godbolt.org/z/1dbYhEavM). this is still 5 instructions because it can combine the load and increment into a single postincrement load, but the return-if-not-equal costs us two instructions instead of one. and the actual speed of that will depend a lot more on things like macro-op fusion and branch prediction. like, it would be easy for the bne to take a lot more than two cycles.
for cortex-m4 gcc also generates the same code for both versions, but of course there's no bxeq on cortex-m4:
1: cmp r3, r1
bne.n 2f
bx lr
2: ldr.w r2, [r3], #4
add r0, r2
b.n 1b
(same in 13.2.0: https://godbolt.org/z/Wc3bv6sWa)so in conclusion risc-v gcc still needs a lot of work to catch up with arm gcc, even to successfully apply broadly-applicable cross-platform optimizations that it successfully applies on arm, which surprised me a lot
postscript: apparently with -O3 instead of -Os gcc does a lot better on risc-v