I’ve been a little shy about using intel software since reading about this years ago
I’ve been a little shy about using intel software since reading about this years ago
Edit, comparison:
$ perf record target/release/gemm-benchmark -d 1024
Threads: 1
Iterations per thread: 1000
Matrix shape: 1024 x 1024
GFLOPS/s: 96.36
$ perf report --stdio -q | head -n3
97.18% gemm-benchmark gemm-benchmark [.] mkl_blas_def_sgemm_kernel_0_zen
1.94% gemm-benchmark gemm-benchmark [.] mkl_blas_def_sgemm_scopy_down16_bdz
0.78% gemm-benchmark gemm-benchmark [.] mkl_blas_def_sgemm_scopy_right4_bdz
After disabling Intel CPU detection: $ perf record target/release/gemm-benchmark -d 1024
Threads: 1
Iterations per thread: 1000
Matrix shape: 1024 x 1024
GFLOPS/s: 129.12
$ perf report --stdio -q | head -n3
97.02% gemm-benchmark libmkl_avx2.so.1 [.] mkl_blas_avx2_sgemm_kernel_0
1.77% gemm-benchmark libmkl_avx2.so.1 [.] mkl_blas_avx2_sgemm_scopy_down24_ea
1.02% gemm-benchmark libmkl_avx2.so.1 [.] mkl_blas_avx2_sgemm_scopy_right4_ea
Benchmarked using https://github.com/danieldk/gemm-benchmark and oneMKL 2021.3.0.I'm really surprised popular numerical computing Python packages don't already have optimized hardware back-ends for things like NumPy... similar to ORC (OIL) which has been around for quite some time:
https://github.com/GStreamer/orc
But I don't know that much about Python under the hood, and I'm willing to be since so many academics work on this there's already optimized FFIs. I've used TensorFlow and it can offload tensor math to GPUs, but only NVIDIA's AFAIK.
I think it is hard to beat modern BLAS implementations for common operations. E.g. Apple Accelerate (which also implement the BLAS/LAPACK APIs) uses undocumented AMX instructions for large speedups compared to an ARM NEON implementation.