The paper would have been better off with more complex examples (e.g. functions that can't be trivially implemented in asm by hand with no register pressure) and fewer comparisons to implementations that could be doing what they are, but aren't.
The paper would have been better off with more complex examples (e.g. functions that can't be trivially implemented in asm by hand with no register pressure) and fewer comparisons to implementations that could be doing what they are, but aren't.
The more complex examples are in Section 5.2; see Figure 8. Granted, we would have liked to have done more, but deadlines are deadlines...
There are some other CPU-related slight inaccuracies in the paper. Prefetching is repeatedly mentioned, even though its effect is negligible when one has a perfectly linear memory access pattern; unaligned loads are mentioned as a performance hit, but they are essentially free on the test processor (2600k, Sandy Bridge).
Matrix multiplication would perhaps be a better example to show the power of clever prefetching.
all that said, it is still very impressive work!
Perhaps somebody with an AVX capable processor can redo the two benchmarks with correct compiler options (i.e. -mavx) and/or with ICC, and replacing the C code of the second benchmark with code that is equivalent to the Haskell code:
double s=0;
for(int i=0; i<n; i++) s += pow(a[i]-b[i],2);I found out gcc supports -fprefetch-loop-arrays, although it is not guaranteed to have a positive effect. In this case, it does appear to run faster than without prefetching. AVX versions also run faster than the standard SSE counterparts even with 128 bit registers. ymm regs are faster than xmm.
gcc is faster than icc actually, and icc runs faster with the plain c version than the vectorized versions. I can't figure out how to make icc prefetch either, the switch that controls it doesn't seem to have an effect.