Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernelsnanduruganesh.github.io·41 pts·rawsh·5