In cases of parallel processing where code executing in the critical path is a significant bottleneck, hand-rolled SIMD can sometimes outperform what gcc or clang spits out by a few clock cycles/iteration. But even then, the sanest way to do that is to let the compiler do most of the legwork and then incrementally tweak its output.
IME gcc is extraordinarily bad at anything but the most basic of auto-vectorizations, but its output from SIMD intrinsics is decent -- not usually quite as good as hand-rolled assembly, but much faster to write (for me anyway, YMMV) and still within 2x performance.