Interesting investigation!
I had an experiment with getting the Rust compiler to vectorise things itself, and it seems LLVM does a pretty good job automatically, e.g. on my computer (x86-64), running `rustc -O bytesum.rs` optimises the core of the addition:
fn inner(x: &[u8]) -> u8 {
let mut s = 0;
for b in x.iter() {
s += *b;
}
s
}
to .LBB0_6:
movdqa %xmm1, %xmm2
movdqa %xmm0, %xmm3
movdqu -16(%rsi), %xmm0
movdqu (%rsi), %xmm1
paddb %xmm3, %xmm0
paddb %xmm2, %xmm1
addq $32, %rsi
addq $-32, %rdi
jne .LBB0_6
I can convince clang to automatically vectorize the inner loop in [1] to equivalent code (by passing -O3), but I can't seem to get GCC to do anything but a byte-by-byte tranversal.[1]: https://github.com/jvns/howcomputer/blob/master/bytesum.c