I don’t think “sparsity” in ML, where like 10...50% is sparse, is the same as sparsity in Physics, where 99.999% (add more 9 with larger problems) of your matrix is sparse.
Half the entries in your matrices need to be 0, then the Hardware will compress them and execute the matrix-matrix multiplication 2x as fast.
On a tensor instruction level.