The problem with autovectorization is that there's no good tooling for it. icc at least has
#pragma simd
to force vectorization, but gcc and clang don't. And there's no way to annotate a loop to say "throw a compile error if this loop fails to autovectorize". You either pray that someone else doesn't cause a silent 4x regression by adding a loop-carried dependency, or you perform some brittle hackery trying to read the output of -fopt-info-vec-missed flag and killing the build when you see a file name + line number that corresponds to a hot loop.At that point, it just becomes easier to use SIMD intrinsics. You're still deferring to the compiler for the instruction scheduling and register allocation, but you no longer have to worry about whether the SIMD is being emitted at all.