- The matrices weren't that big. Numpy has cffi overhead.
- The perf difference was much more noticeable with _peak_ throughput rather than _mean_ throughput, which matters for almost no applications (a few, admittedly, but even where "peak" is close to the right measure you usually want something like the mean of the top-k results or the proportion with under some latency, ...).
- The benchmarking code they displayed runs through Python's allocator for numpy and is suggestive of not going through any allocator for the C implementation. Everything might be fine, but that'd be the first place I checked for microbenchmarking errors or discrepancies (most numpy routines allow in-place operations; given that that's known to be a bottleneck in some applications of numpy, I'd be tempted to explicitly examine benchmarks for in-place versions of both).
- Numpy has some bounds checking and error handling code which runs regardless of the underlying implementation. That's part of why it's so bleedingly slow for small matrices compared to even vanilla Python lists (they tested bigger matrices too, so this isn't the only effect, but I'll mention it anyway). It's hard to make something faster when you add a few thousand cycles of pure overhead.
- This was a very principled approach to saturating the relevant caches. It's "obvious" in some sense, but clear engineering improvements are worth highlighting in discussions like this, in the sense that OpenBLAS, even with many man-hours, likely hasn't thought of everything.
And so on. Anyone can rattle off differences. A proper explanation requires an actual deep-dive into both chunks of code.
What's really surprising is that many modern languages are still depending on OpenBLAS for examples Matlab, Julia, Mojo, etc. But to be fair they probably have their own reasons, right?
[1] Numeric age for D: Mir GLAS is faster than OpenBLAS and Eigen (2016):
http://blog.mir.dlang.io/glas/benchmark/openblas/2016/09/23/...
[2] Vastly outperforming LAPACK with C++ metaprogramming (2018):
https://wordsandbuttons.online/vastly_outperforming_lapack_w...
[3] Outperforming LAPACK with C metaprogramming (2018):
https://wordsandbuttons.online/outperforming_lapack_with_c_m...
https://en.wikipedia.org/wiki/X86-64#Microarchitecture_level...
E.g. https://github.com/numpy/numpy/blob/main/numpy/_core/src/com...
Could be MKL (i believe the conda version comes with it) but it could also be an ancient version of OpenBLAS you already had installed. So yeah, being faster than np.matmul probably just means your NumPy is not installed optimally.