- a refresher on differential equations
- legendre polynomials
- state spaced models; you need to grok the essence of
x' = Ax + Bu
y = Cx
- discretization of S4
- HiPPO matrix
- GPU architecture (SRAM, HBM)
Basically, transformers is an architecture that uses attention. Mamba is the same architecture that replaces attention with S4 - but this S4 is modified to overcome its shortcomings, allowing it to act like a CNN during training and an RNN during inference.
I found this video very helpful: https://www.youtube.com/watch?v=8Q_tqwpTpVU
His other videos are really good too.
https://srush.github.io/annotated-s4/
The matrices that make up the state space (A, B and C) are constant in S4. This allowed them to represent some of the math operations as a convolution (which can be parallelized).
The difference between S4 and Mamba is that these matrices are input-dependent in Mamba. Plus they add in some CUDA stuff ("parallel scan") to make it faster to compute on a GPU even if these matrices are not constant.
Yannic Kilcher's video on Mamba might also be a good resource: https://youtu.be/9dSkvxS2EB0
https://blog.oxen.ai/mamba-linear-time-sequence-modeling-wit...
I'm still not convinced on Mamba's performance on Natural Language tasks, but maybe it's just because they haven't trained a large enough model on enough data yet.
Feel free to join here: https://lu.ma/oxenbookclub
Here is my intuition about Mamba. It’s a linear discrete time signal filter where the filter blocks are conditioned on the input at each time step as well as receiving the input as a normal filter would. Think of the predecessor to Mamba as a discrete time filter without this conditioning. That’s the first innovation.
The other innovation is some lovely algebra allowing them to reorganize things so they can be computed much more efficiently on current GPU hardware, with respect to slow and fast memory.
State space models are very simple conceptually. It’s kind of amazing that Mamba does so well on long sequence modeling because the architecture is ridiculously simple. But when you consider that the state matrices are conditioned on the input, it all makes sense. Somehow, they learn how to adapt the “filter” based on the input. That’s where the intelligence lies.