Why aren't they showing benchmarks that compare it to other BLAS implementations? How does it compare to the GEMM in atlas, cblas, intel's mks, GOTOBlas, or any other library that implements GEMM? Is writing 'jitted' asm like this better than writing fortran with -march=native?