Here's some data:
http://lolengine.net/attachment/blog/2011/9/17/playing-with-...
Intel(R) Xeon(R) CPU E5-1620 0 @ 3.60GHz stepping 07
uname -a: Linux 3.5.0-41-generic #64-Ubuntu SMP
clang++ -v: clang version 3.2
icpc -v: icpc version 14.0.1
g++ -v: gcc version 4.7.2:
Compiled with "-std=c++11" plus the options mentioned in the names. At a glance, all versions produced the same results to reasonable precision (no gross errors). "-ffast-math" only worked with g++, and did not offer substantial improvement, so I've omitted it from the tests. g++_O3 icpc_O3 clang++_O3
sin: 18.193 ns 7.0174 ns 14.6862 ns
sin1: 22.5272 ns 8.0225 ns 9.6777 ns
sin2: 14.9128 ns 4.7247 ns 7.1016 ns
sin3: 18.943 ns 5.1121 ns 6.169 ns
sin4: 13.9225 ns 4.8979 ns 5.6666 ns
sin5: 14.2042 ns 5.0435 ns 6.3257 ns
sin6: 12.2543 ns 4.2955 ns 5.2681 ns
sin7: 11.6969 ns 5.0538 ns 5.5793 ns
So yes, these optimizations can still have a positive effect, but generally less than that of using a better compiler. In this case, that means if you care about performance use Intel, or if not possible, use CLang.But what about vectorization? Perhaps the compiler can do a better job if you let it make full use of modern SIMD:
g++_O3_mavx icpc_O3_mavx clang++_O3_mavx
sin: 18.1148 ns 3.7879 ns 14.1987 ns
sin1: 22.3675 ns 6.4376 ns 9.2589 ns
sin2: 14.5665 ns 3.2823 ns 6.2089 ns
sin3: 18.3841 ns 3.6652 ns 5.2409 ns
sin4: 13.0917 ns 3.0374 ns 4.9839 ns
sin5: 13.1144 ns 3.598 ns 5.177 ns
sin6: 11.7605 ns 2.8766 ns 4.2239 ns
sin7: 11.4222 ns 3.3261 ns 4.612 ns
Yes, it looks like Intel gains quite a bit from vectorizing the result with 256-bit AVX (4-wide for doubles). If you care about autovectorization, use Intel.This performance here generally matched my expectations. I have no affiliation with Intel other than having a free academic license, but find that in general their compiler offers better performance than GCC or CLang. In this case, Intel's best (sin6 with AVX) is 4x faster than GCC's best. I'd usually expect something more like a 20%-40% improvement, but vectorization is one of Intel's strong suits.
For those who wish to explore the reasons for the difference in performance, here are the results of 'objdump -C -d' for the functions in question:
g++_O3: http://pastebin.com/VstqvcHJ
icpc_O3: http://pastebin.com/3LjtXAhS
clang++_O3: http://pastebin.com/aSyCyULh
g++_O3_mavx: http://pastebin.com/81vYueEp
icpc_O3_mavx: http://pastebin.com/XRdCuusV
clang++_O3_mavx: http://pastebin.com/KJqUpUBE
Personally, I'd be very interested to see what an experienced x64 SIMD programmer could do to improve these further. My usual estimate is a further 2x, but I don't know how well that applies to this case. I'd also welcome analysis of the code produced by the compilers.