Here are some performances with higher optimization parameters and with the Intel C Compiler. All run with a Haswell Xeon (E5-2698 v3). Using icc, for the larger problem sizes you get performance speedup that very similar to the speedup from the hard-coded asm.
With same arguments as blog post:
gcc -O1 -o fft-test fft-test.c fft-portable.c -lm:
...
262144 min=15721397 mean=15748901 sd=0.12%
524288 min=34474976 mean=34546878 sd=0.18%
1048576 min=99814826 mean=100391926 sd=0.26%
...Adding -march=native and -O2:
...
262144 min=14668178 mean=14701671 sd=0.14%
524288 min=31297081 mean=31395468 sd=0.14%
1048576 min=94027867 mean=94196746 sd=0.13%
...With Intel C Compiler (-O2 -xHost):
...
262144 min=14258167 mean=14297606 sd=0.16%
524288 min=31126755 mean=31212153 sd=0.23%
1048576 min=92893600 mean=93491419 sd=0.46%
...Can provide more numbers if people want them. EDIT: looks like I was using the less optimized C code, will re-run with the better performing C code when I get the chance.