First we can use Laser, which was my initial BLAS experiment in 2019. At the time in particular, OpenBLAS didn't properly use the AVX512 VPUs. (See thread in BLIS
https://github.com/flame/blis/issues/352 ), It has made progress since then, still, on my current laptop perf is in the same range
Reproduction:
- Assuming x86 and preferably Linux.
- Install Nim
- Install a C compiler with OpenMP support (not the default MacOS Clang)
- Install git
The repo submodules MKLDNN (now Intel oneDNN) to bench vs Intel JIT Compiler
```
git clone https://github.com/mratsim/laser
cd laser
git submodule update --init --recursive
nim cpp -r --outdir:build -d:danger -d:openmp benchmarks/gemm/gemm_bench_float32.nim
```
This should output something like this
```
Laser production implementation
Collected 10 samples in 0.230 seconds
Average time: 22.684 ms
Stddev time: 0.596 ms
Min time: 21.769 ms
Max time: 23.603 ms
Perf: 624.037 GFLOP/s
OpenBLAS benchmark
Collected 10 samples in 0.216 seconds
Average time: 21.340 ms
Stddev time: 3.334 ms
Min time: 19.346 ms
Max time: 27.502 ms
Perf: 663.359 GFLOP/s
MKL-DNN JIT AVX512 benchmark
Collected 10 samples in 0.201 seconds
Average time: 19.775 ms
Stddev time: 8.262 ms
Min time: 15.625 ms
Max time: 43.237 ms
Perf: 715.855 GFLOP/s
```
Note: the Theoretical peak limit is hardcoded and used my previous machine i9-9980XE.
It maybe that your BLAS library is not named libopenblas.so, you can change that here: https://github.com/mratsim/laser/blob/master/benchmarks/thir...
Implementation is in this folder: https://github.com/mratsim/laser/tree/master/laser/primitive...
in particular, tiling, cache and register optimization: https://github.com/mratsim/laser/blob/master/laser/primitive...
AVX512 code generator: https://github.com/mratsim/laser/blob/master/laser/primitive...
And generic Scalar/SSE/AVX/AVX2/AVX512 microkernel generator (this is Nim macros to generate code at compile-time): https://github.com/mratsim/laser/blob/master/laser/primitive...
I'll come back later with details on how to use my custom HPC threadpool Weave instead of OpenMP (https://github.com/mratsim/weave/tree/master/benchmarks/matm...). As a side bonus it also has parallel nqueens implemented.