Mamba-2 – State Space Duality
tridao.me
tridao.me
Theoretical stuff aside, Mamba-2's performance seems to scale slightly better than original Mamba: https://tridao.me/assets/img/2024-05-31-mamba-2/pile_8k_mamb...
Here's the code implementing Mamba-2: https://github.com/state-spaces/mamba/blob/main/mamba_ssm/mo...
Great work by Tri Dao (of FlashAttention fame) and Albert Gu, as usual.
The key question, for me and many others, is whether Mamba, Mamba-2, RWKV, and other linear RNN / linear attention models will ever match the performance of standard Softmax attention. My understanding and experience is that all the linear attention models out there [b] still underperform Softmax attention on things like recall tasks.[c]
---
[a] https://arxiv.org/abs/2006.16236
[b] https://github.com/topics/linear-attention / https://github.com/topics/linear-attention-model -- this list is by no means complete!
Source? I don't think this is remotely true based on even the most recent benchmarks I've seen.
> precisely recall things from it on demand - "precisely" is load-bearing here, and the question at hand. (no, LLMs don't have perfect or precise recall, even from in-context material, and it's not particularly close)
> translate Farsi into Hawaiian and then speak it - 0 models, is this supposed to be a gpt-4o reference? What does this have to do with "perfect recall?" Let's pretend it does. Are you claiming that gpt-4o has a 100% success rate translating Farsi to Hawaiian?
> Can you give instant, and nearly always correct to almost any textbook question - sobs in just spent 4 days benchmarking. For the record, best non-RAG I have is gpt-4o at 86% on MedQA and LegalBench. With RAG, 93%. RAG gpt-4o just barely scrapes by my cofounders USMLE scores. It's excellent. But it's not superhuman.
> All of that is recall - no, translation is not recall
Thank you. Yes, that's obviously true... but keep in mind that human beings can pull up information from notes, books, URLs, etc. and store that information in "short-term memory" at any time, as needed, for cognitive tasks. To match or exceed peak human performance, AI models must be able to do that too -- whether it's via a long context window or some other mechanism (e.g., searching on external storage) is a secondary issue. I'd say recall is a prerequisite to match peak human performance.
a) a linear SSM (a form of RNN?) is equivalent to Attention without the scaling and softmax; and
b) Attention is "all you need" and the thing that made Transformers radically outperform all the previous architectures like LSTMs that used to dominate NLP;
does that imply c) the scaling and softmax parts of the attention equation, in particular, is the magic touch that makes Transformers work so well?
An important role is held by the softmax function which normalizes the attention scores, allowing the model to weigh different parts of the input sequence dynamically. This means that, unlike RNNs which sequentially process inputs and update states, Transformers can directly access and prioritize information from any part of the sequence, and they are not slower for T < N.
---
the actual flops involved are similar to the original SSM-based version, but that's harder to formulate as strictly matrix multiplications