> while keeping a constant state size for generating the output
Isn't Mamba's attention mechanism linear relative to the input? The innovation as I understand it is that attention in Mamba *isn't* quadratic (like transformers).
Isn't Mamba's attention mechanism linear relative to the input? The innovation as I understand it is that attention in Mamba *isn't* quadratic (like transformers).
Both statements are correct
Mamba adds a dependence on the inputs that makes language modeling competitive with transformers, but that prevents using the fft approach. So they switch to a method using parallel prefix scan.