> Citation needed. Last time I checked LLVM beat GCC on -O3.
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
https://www.phoronix.com/review/gcc-clang-eoy2023/8
There is tons more, just ask your favourite internet search engine.
> Why should two nearly idenfical operations have such wildly different performance? And why isn't the safer/saner interface more performant?
Because in hardware, checking a zero-flag on a register is very cheap. Big NAND over all bits of a sub-register, big OR over all those flags for the whole SIMD register. Very short path, very few transistors.
Checking a length is expensive: Increment a length register through addition, including a carry, which is a long path. Compare with another register to out if you are at the last iteration. Compare usually is also expensive since it is subtract with carry internally, even though you could get away with xor + zero flag. But you can't get around the length register addition.
There can be an optimisation because you do have to increment the address you fetch from in any case. But in most modern CPUs, that happens in special pointer registers internally, if you do additional arithmetics on those it'll be expensive again. x86 actually has more complex addressing modes that handle those, but compilers never really used them, so those were dropped for SIMD.