This was my function
for (auto v : array) {
if (v != 0)
return false;
}
return true;
clang emits basically the same thing yours does. But gcc ends up just really struggling to vectorize for large numbers of array.Here's gcc for 42 elements:
https://godbolt.org/z/sjz7xd8Gs
and here's clang for 42 elements:
https://godbolt.org/z/frvbhrnEK
Very bizarre. Clang pretty readily sees that it can use SIMD instructions and really optimizes this while GCC really struggles to want to use it. I've even seen strange output where GCC will emit SIMD instructions for the first loop and then falls back on regular x86 compares for the rest.
Edit: Actually, it looks like for large enough array sizes, it flips. At 256 elements, gcc ends up emitting simd instructions while clang does pure x86. So strange.
GCC trunk seems to like using `bool` so we may eventually be able to retire the hack of using `int`.
This is purely me looking at the emitted assembly and being surprised at when the compilers decide to deploy it and not deploy it. It may be the case that the SIMD instructions are in fact slower even though they should theoretically end up faster.
Both compilers are simply using heuristics to determine when it's fruitful to deploy SIMD instructions.