the benchmarks are not about GEMM, but real-world deep learning workloads which could have very different characteristics from GEMM
The GEMM example was just there as the details of the optimization have been published, unlike most other hand-tuned assembler routines for DNN workloads.