Mamba: New SSM arch with linear-time scaling that outperforms Transformers
github.com
github.com
However the greater (5x) inference bandwidth makes this super appealing especially for democratizing AI and enabling the GPU poor. This could very well be a watershed moment for SSMs, similar to how Transformers boasted improvements to training and inference speed in 2019.