I suspect this comment is at least partially directed at me. Did you read the paper? Excluding some O(1) operations, the goal of the second benchmark is this:
double s = 0;
for(int i=0; i<n; i++) s += pow(a[i]-b[i],2);
return s;
But they did not use this C code. The C code they used made 3 calls to the BLAS library. The end result is that the C code is doing something equivalent to this: double x=0, y=0, z=0;
for(int i=0; i<n; i++) x += a[i]*a[i];
for(int i=0; i<n; i++) y += a[i]*b[i];
for(int i=0; i<n; i++) z += b[i]*b[i];
return x - 2*y + z;
While this does return the same result (assuming infinite precision arithmetic) it is obviously not the way anybody would do it, since it's doing 3 times the work. Even worse, because their code is calling into the BLAS library for each loop, the compiler is explicitly prevented from optimizing the three loops and memory accesses by combining them into one. Note that the Haskell code is doing the efficient single loop. So yes, it is poorly written code.