An Empirical Study of Mamba-Based Language Models
arxiv.org
arxiv.org
Seems like it scales better than transformers, but this would only be really obvious at parameter counts far in excess of the experiments in this paper.
The rate of improvements seems quick enough right now that if you started training a huge model now that you might regret spending all that money on an architecture that is obsolete by the time you are finished.
That said, if you keep waiting you never get around to it.
It would be nice to see a large parameter mamba family data point though.