-ffast-math allows the multiply-adds to be reassociated, enabling much better approaches. Clang with -O2 -ffast-math produces good code (vfmadd132ps with 4 independent accumulators), I can't get GCC to produce good code with any flags.
`-funroll-loops` by default unrolls 8 times. Unrolling beyond the number of accumulators is wasteful for simple operations like dot procuts or summations, so you may want to control that with `--param -max-unroll-times=4`.
Thus, the following works and will produce generally faster code than LLVM (because LLVM doesn't vectorize the remainder, giving you potentially large numbers of scalar operations):
`-Ofast -funroll-loops --param max-unroll-times=4 -fvariable-expansion-in-unroller --param max-variable-expansions-in-unroller=4`
Godbolt: https://godbolt.org/z/4PXSqs
$ gfortran-10 -c dot.f90 -Ofast -fopt-info
dot.f90:3:7: optimized: loop vectorized using 16 byte vectors
dot.f90:3:7: optimized: loop with 2 iterations completely unrolled (header execution count 64530389)
$ gfortran-10 -c dot.f90 -Ofast -fopt-info -funroll-loops --param max-unroll-times=4 -fvariable-expansion-in-unroller --param max-variable-expansions-in-unroller=4
dot.f90:3:7: optimized: loop vectorized using 16 byte vectors
dot.f90:3:7: optimized: loop with 2 iterations completely unrolled (header execution count 64530389)
dot.f90:4:0: optimized: loop unrolled 3 times