"skipping Nvidia’s closed-source CUDA libraries, such as cuBLAS, in favor of open-source libraries, such as cutlass"
Eww, so it'll be 90% slower?
cuBLAS is insanely optimized. It's the reason why CUDA dominates the field even though one could in theory use a ROCm build of TensorFlow or PyTorch. But performance is abysmal without cuBLAS. So the entire article appears to hinge on a huge assumption: What if Free Open Source was equivalent to NVIDIA's top-tier priorietary library? (It's not)