I benchmarked compilers for some of my stuff just over a year ago (GCC, LLVM, ICC): http://imagine-rt.blogspot.co.nz/2014/12/c-compiler-benchmar...
and surprisingly I found ICC was no longer always the fastest, whereas two years previously in tests I'd regularly done against GCC 4.1 and 4.4 it would always win, sometimes by a factor of 2x.
Also, the article brought up the idea of block sizing, which sounds compelling. But the writer failed to produce a benchmark for it which did better than baseline on ICC, and then kept on writing as if block sizing had merit without even commenting on this discrepancy.
For matrices of size 4000x4000 there was significant benefit for blocking the loops - please check the table again. This benefit continues as matrices continue to get larger. I specifically used two different size matrices as examples so that the reader could see that blocking is not a panacea or cure all. So as you point out for the smaller case there was no benefit - I wanted the reader to see that, so I included both data sets. The value of block size varies depending on the matrix size and the cache size of the system you are running on. When the data sets or number of iterations you are running on is relatively small (compared to cache size) the blocking will not provide benefit; don't blindly apply it. The matrix multiply was selected for pedagogical purposes - it is easy to illustrate and easy to explain. I was limited to article length - see the link to the article on finite difference methods for more on value of blocking. If you are really interested in blocking for matrix multiply there many other good articles that go into this in much more detail (and the fastest ones do block into submatrices and perform additional optimizations). (Please note the optimal blocking may be rectangular non square blocks - I did not code the general case.)