> This is problematic because not only do we get an extra read
I don't understand why you're speculating on this. The author provided the assembly that only has 1 read.
FWIW, I just tested this on my machine. I get the following results:
Intel Core i7-7800 3.5Ghz / Ubuntu 20 / g++ 9.3.0 Rolled loop : 2.08 Gflops | 16.7 GB/s bandwidth 8-unrolled loop : 2.48 Gflops | 19.8 GB/s bandwidth
Here's the unrolled assembly (8 iters), with -O3:
.L3:
addl (%rdi), %eax
addq $32, %rdi
movl %eax, -32(%rdi)
addl -28(%rdi), %eax
movl %eax, -28(%rdi)
addl -24(%rdi), %eax
movl %eax, -24(%rdi)
addl -20(%rdi), %eax
movl %eax, -20(%rdi)
addl -16(%rdi), %eax
movl %eax, -16(%rdi)
addl -12(%rdi), %eax
movl %eax, -12(%rdi)
addl -8(%rdi), %eax
movl %eax, -8(%rdi)
addl -4(%rdi), %eax
movl %eax, -4(%rdi)
cmpq %rdi, %rdx
jne .L3