Vision Mamba: Efficient Visual Representation Learning with Bidirectional SSM
arxiv.org
arxiv.org
Isn't Mamba's attention mechanism linear relative to the input? The innovation as I understand it is that attention in Mamba *isn't* quadratic (like transformers).
Both statements are correct
Mamba adds a dependence on the inputs that makes language modeling competitive with transformers, but that prevents using the fft approach. So they switch to a method using parallel prefix scan.
The latest iteration has promising results for language but no one has yet trained and released a big (7B+) model to see how it scales.
Really?! Come-on.