Roofline plots [1] are framework to visualize system design from this perspective.
Roofline plots [1] are framework to visualize system design from this perspective.
In this article the deciding factor seems to be the startup cost because the application has placed the data on the CPU memory side, and is considering shipping out to GPU memory just for this computation.
I don’t think it’s accurate that only trivial implementations use the direct o(n^3) algorithm. AFAIK high performance BLAS implementations just use highly optimized versions of it.
I just wish I understood the tricks done to make it so fast so I could implement my own for variations for which there are no pre-existing BLAS implementations. The best BLAS implementations are all closed source sadly.
NVidia open-sourced CUTLASS [0] some years ago and it achieves pretty competitive performance compared to e.g. the closed-source cuBLAS.
Keen observers will notice that Strassen is not used in CUTLASS.
CUDNN supports FFT to do matmul as well as convolution/correlation and can also be configured to automatically use the best algorithm.
In some cases the FFT method has the incidental side-benefit of data reuse, like in the case of FIR filters, where the data allows for partitioned convolution.