Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels
nanduruganesh.github.io
nanduruganesh.github.io
Has anyone used the new Minimax M3 model? I’m curious how it compares with Deepseek V4 and GLM 5.2 and other larger open weights models.
[Update: their cheapest token plan has been removed, I guess its back to GLM now]
Right now M3 is not far behind DS4, but I belive DS4 will improve much more with each round of training. It simply has a bigger brain, it just needs to fill it with more information.
Such lazy, much farming
https://github.com/fla-org/native-sparse-attention?utm_sourc...