It generates vector instructions for all k8-like architectures, with wider instructions starting from -march=sandybridge.
vmovss, vmulss and vaddss instructions are not vector. They only load/store, multiply or add the lowest lane of their operands. The equivalent vector instructions are vmovaps, vmulps and vaddps respectively, these would load, multiply or add complete vectors, not just the lowest lanes of them.
-ffast-math allows the multiply-adds to be reassociated, enabling much better approaches. Clang with -O2 -ffast-math produces good code (vfmadd132ps with 4 independent accumulators), I can't get GCC to produce good code with any flags.
`-funroll-loops` by default unrolls 8 times. Unrolling beyond the number of accumulators is wasteful for simple operations like dot procuts or summations, so you may want to control that with `--param -max-unroll-times=4`.
Thus, the following works and will produce generally faster code than LLVM (because LLVM doesn't vectorize the remainder, giving you potentially large numbers of scalar operations):
`-Ofast -funroll-loops --param max-unroll-times=4 -fvariable-expansion-in-unroller --param max-variable-expansions-in-unroller=4`
Godbolt: https://godbolt.org/z/4PXSqs
$ gfortran-10 -c dot.f90 -Ofast -fopt-info
dot.f90:3:7: optimized: loop vectorized using 16 byte vectors
dot.f90:3:7: optimized: loop with 2 iterations completely unrolled (header execution count 64530389)
$ gfortran-10 -c dot.f90 -Ofast -fopt-info -funroll-loops --param max-unroll-times=4 -fvariable-expansion-in-unroller --param max-variable-expansions-in-unroller=4
dot.f90:3:7: optimized: loop vectorized using 16 byte vectors
dot.f90:3:7: optimized: loop with 2 iterations completely unrolled (header execution count 64530389)
dot.f90:4:0: optimized: loop unrolled 3 times