[0] https://en.m.wikipedia.org/wiki/Basic_Linear_Algebra_Subprog...
[0] https://en.m.wikipedia.org/wiki/Basic_Linear_Algebra_Subprog...
Do people go further than using instrinsics for let's say AVX?
Portions of the algorithm have been translated into assembly for ARM and x86. Shaving even a couple percent off something like motion compensation search will add up to meaningful gains. See also the current reference implementation of JPEG: https://github.com/libjpeg-turbo/libjpeg-turbo/tree/main/sim...
[0] https://github.com/xianyi/OpenBLAS/issues/1968
[1] https://github.com/xianyi/OpenBLAS/blob/develop/kernel/x86_6...
[2] https://github.com/xianyi/OpenBLAS/blob/23693f09a26ffd8b60eb...
Why wouldn’t it be? Compilers haven’t advanced tremendously in the past two decades in terms of optimizations and don’t have much new to add to high performance SIMD numeric kernels.