If you want to really dive into the math and motivation for this entirely new class of models (state space models) I highly recommend reading Albert Gu’s thesis.
https://searchworks.stanford.edu/view/14784021
I tried to read the Mamba paper and I was lost but after reading his thesis it was a lot to understand it.
The part I’m still struggling with is the section on how Mamba trains so efficiently, it’s surprising since it is actually autoregressive and step by step and I thought this is why LSTMs were dropped