Wow, even at similar optimization levels? -O3 for each?
Wow, even at similar optimization levels? -O3 for each?
gfortran: -O2 -ftree-vectorize -funroll-loops ifort: -O3
I don't have the timing results anymore. This was in 2018 on Xeon Platinum 8168.
I recently tried replacing -O2 with -Ofast -ffast-math to gfortran settings, which gives about 18% speed up. So still far from Intel. I recently proposed it here [1].
Coming up with these transformations is not the difficult part. That would be actually implementing these transformations correctly and keep the whole compiler optimization framework maintainable.
And of course there is certainly a cost for the improved runtime. Primarily in compile time and code size. For the Intel compiler it generally makes sense to sacrifice these in favor of runtime performance, because it is used on these kind of scientific codes a lot, but for gfortran the balance might different.