That would imply that the C version of the C++ code would be just as fast, if only it were not "poorly written".
However there are some types of code (including this kind) where having C++ available to generate the code for you makes it almost inherently faster than C.
You could certainly reimplement specific matrix operations for very specific types in C, but the C++ Eigen library does this all for you at compile time, and does it the best way for each combination of types and operations.
If you do this in C, even with macros everywhere, it's going to be much more difficult, especially since you can't (AFAIK) declare and discard temporary types within macros, though C11 at least allows you to refer generically to types within macros. The canonical example for this kind of disparity is C's qsort() vs. C++ templated sorts.
This is kind of the point to the paper. Performance of this sort is only possible in C++ if you're careful to follow the instructions provided with Eigen about using templated functions, and some very smart devs had to write Eigen in the first place. The Haskell code comes fairly close starting from the source code you'd write naïvely, and without tons of optimization on the stream fusion code generator.
The paper would have been better off with more complex examples (e.g. functions that can't be trivially implemented in asm by hand with no register pressure) and fewer comparisons to implementations that could be doing what they are, but aren't.
The more complex examples are in Section 5.2; see Figure 8. Granted, we would have liked to have done more, but deadlines are deadlines...
There are some other CPU-related slight inaccuracies in the paper. Prefetching is repeatedly mentioned, even though its effect is negligible when one has a perfectly linear memory access pattern; unaligned loads are mentioned as a performance hit, but they are essentially free on the test processor (2600k, Sandy Bridge).
Matrix multiplication would perhaps be a better example to show the power of clever prefetching.
all that said, it is still very impressive work!
Perhaps somebody with an AVX capable processor can redo the two benchmarks with correct compiler options (i.e. -mavx) and/or with ICC, and replacing the C code of the second benchmark with code that is equivalent to the Haskell code:
double s=0;
for(int i=0; i<n; i++) s += pow(a[i]-b[i],2);I found out gcc supports -fprefetch-loop-arrays, although it is not guaranteed to have a positive effect. In this case, it does appear to run faster than without prefetching. AVX versions also run faster than the standard SSE counterparts even with 128 bit registers. ymm regs are faster than xmm.
gcc is faster than icc actually, and icc runs faster with the plain c version than the vectorized versions. I can't figure out how to make icc prefetch either, the switch that controls it doesn't seem to have an effect.