1. SSMs are a type of recurrent model that can scale linearly with sequence length, making them more efficient than Transformers. However, prior SSMs struggled with discrete sequence modeling tasks like language.
2. This paper augments SSMs with a "selection mechanism" that allows the model dynamics to depend on the input, giving it the ability to selectively remember or forget information. This makes SSMs effective on tasks requiring discrete reasoning.
3. They design an efficient parallel scan algorithm to implement selective SSMs on GPUs. Despite recurrency, this achieves up to 5x higher throughput over Transformers in benchmarks.
4. They simplify prior SSM architectures into a new model called Mamba. On language modeling, Mamba matches or exceeds Transformers of 2-3x its size, while retaining linear scaling. It also achieves state-of-the-art results on audio, genomics, and synthetic tasks requiring long-term reasoning.
This work makes SSMs truly competitive with Transformers through selectivity and efficient engineering. Mamba matches or beats Transformers on major modalities while being substantially more efficient in computation, memory, and scaling to long sequences. If replicated, it's arguably the first linear-time architecture with Transformer-quality performance!!
Can't wait to see the code!