FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
github.com
github.com
One downside is that as hardware/software becomes more specialized to one architecture, it becomes more difficult to explore very different architectures.
It's a bit surprising that just a longer range transformer architecture enabled by better use of the GPU was the solution though.