xLSTM: Extended Long Short-Term Memory
arxiv.org
arxiv.org
In the scaling law comparison, I wonder if it is reasonable to compare number of parameters between Llama, Mamba, RWKV, xLSTM? Isn't compute time more relevant? E.g. in the figure about scaling laws, replace num of params by compute time.
Specifically, the sLSTM has still recurrence (memory mixing) in it, i.e. you cannot fully parallelize the computation. So scaling up Transformer could still look better when you look at compute time.
It seems neither the code nor the model params are released. I wonder if that will follow.
For really large models, it is in fact easier to achieve peak flops because computation required scales faster than memory bandwidth required(square vs cube).
> Medium sized transformer models are generally not trained with sequence parallelism, but sequence parallelism is getting more common with transformer training
Is there some word missing? You mean it's more common for large-sized Transformers?
> computation required scales faster than memory bandwidth required (square vs cube)
That is an interesting thought. I'm trying to understand what exactly you mean. You mean, computation time is in O(N^2) where N is the sequence length, while required memory bandwidth is in O(N^3)? Why is that?
In terms of difficulty of implementation it's arguably much easier than pipeline parallelism, which I'd argue is the hardest kind (at least to implement it efficiently without bubbles), and takes the most lines of code to implement (especially in Jax, where sequence parallelism is almost trivial).
As a clarification: The speed for training will be on par with FlashAttention-2, when fully optimized and only including the mLSTM. For decoding/inference both are very close to Mamba as xLSTM is a recurrent architecture. The sLSTM has memory mixing, that is state tracking capabilities, for problems Transformers and State Space Models (and any other sequence-parallelizable architecture) cannot solve fundamentally.
But you would want to include sLSTM as well to get the best performance, right? How does the speed compares in that case? Specifically when scaling up.
Surely whether a big model using a certain system exists is only a matter of the choices of those with sufficient resources to train it. That's only a matter of their beliefs, not about actual model performance.
Unless you give them chain of thought. In which case they do great.
Can you summarise how the model in your paper differs from this implementation of xLSTM ?
Can you opine on how the model will fare on hardware that is optimized for transformers? There is so much investment in accelerating the transformer arch[1][2], will xLSTM / sLSTM benefit as well, or will the hardware optimizations give transformers enough of an advantage that it’s hard to compete on general purpose hardware?
2. https://www.embedded.com/ai-chip-features-hardware-support-f...
Can you explain this statement more if you have time? Are you saying the recurrent architecture of xLSTM enables fast inference on par with Mamba? Or the xLSTM architecture slows it down so that its inference is as slow as mamba?
So in mLSTM, each unit of the vector c is now a matrix (so a 3d tensor)? And we refer to each matrix as a head?
Having a bit of issue understanding this fundamental part
For the matrix 'C' state, there are also heads/cells in that sense that you have multiple, but they don't talk to each other. So yes, you can view that as a 3D tensor. And here, the matrix is the fundamental building block / concept.
If you mean that you cannot fully parallelize inference, this might be true but also not quite relevant since the computational demands of inference are low. And you can always "parallelize" training to some extent, just by training larger batches.
https://betterexplained.com/articles/colorized-math-equation...
The claim is something than will replace the transformer, a technology powering a good chunk of AI companies.
The paper's authors seems to be either from a public university, or Sepp Hochreiter's private company or labs nx-ai.com https://www.nx-ai.com/en/xlstm
Where is the code ? What is the license ? How are they earning money ? Why publish their secret recipe ? Will they not be replicated ? How will the rewards be commensurate with the value their algorithm bring ? Who will get money from this new technology ?
At least capitalists have something to fight over that’s worth fighting for (money). Academics will bitterly fight over the dumbest, least important shit. There’s a law about how the less something matters, the more political the fights over it will be.
The authors are positioning themselves as a company and not merely academics :
A Sepp Hochreiter's video from 6 months ago hyping xLSTM :
https://youtu.be/hwIt7ezy6t8?feature=shared&t=561 in which he state his intent to raise €300M to make a european alternative to openai's GPT for niche domains thanks to this new method that will allow to train for cheaper and better.
He recently received (2023) € 35,000 in prize money at the 5th annual German AI Award.
https://www.jku.at/en/festival-university/media/detail/news/...
Or is it just an academic tactic to get more funding ? To extract more work from PhD students by making them think they are going to strike it big ?
How are they intending to build a moat if they publish their papers ? Will this technology be encumbered by patents/license ?
In this specific xLSTM case, the industry has matured, they are just one among many (Mamba, S3Ms, transformers-variants... ), they have already been sitting on it for at least 6 months, I don't see what their play is.
An other case study that's probably interesting, are the authors of the Adam Paper, https://arxiv.org/abs/1412.6980 , (Awarded "2020: The Adam optimization paper is the world's #1 most cited scientific paper of the past five years"). Probably a few (10?,100?) billions worth value created. You can find the authors bios http://dpkingma.com/ https://jimmylba.github.io/
I think there is a huge problem with the capture and sharing of value in the whole deep-learning industry. Academia's naivety plays a role in it, Generational Shift technologies are badly rewarded. Incremental Shift technologies aren't rewarded at all.
Powerful technologies into many hands with low rewards for their creators while the value they generate keeps going to the same pockets. That's a recipe for disaster.
Will be a fun thing to come back in a few years to see how it had unfold.
LSTMs where very popular for a while (I think the first good version of Google Translate used them) but they had two critical downsides: their performance went down with longer outputs, and they where a bit annoying to parallelize because computing the output for the 10th word required first computing the output of the previous 9 words - no way to use 10 parallel computers. The first problem was solved with Attention, a scaffolding method that prevented degradation over longer sequences. Eventually someone realized that Attention was doing most of the heavy lifting, built an attention-only network that could be easily parallelized (the Transformer), and LSTMs lost the top place.
Are xLSTMs better? On paper I'd say they could be - they seem to have a solid theory and good results. Will they dethrone Transformers? My guess is no, as it wouldn't be the first time that the "better" technology ends up losing against whatever is popular. Having said that, it is entirely possible that some inherently recurrent tasks like stock price prediction could get a boost from this technology and they may find their place.
So GPT-3 Medium (from the GPT-3 paper) - feels pretty disingenuous to list that as no one is referencing that model when they say "GPT-3", but the 175B model.
I wasn't aware that size of the model (356M) was released- what am I missing here?
I also think it's relatively well understood that (with our current methods) transformers have a tipping point with parameter count, and I don't know of any models less than ~3B that are useful- arguably 7B.
Compare these benchmarks to, say, the RWKV 5/6 paper https://arxiv.org/abs/2404.05892
You should see the work on ReFT coming from mannings group showing that you can instruction fine tune models by modifying like, 0.00001% of the parameters. By doing it this way, you significantly mitigate the risk of catastrophic forgetting.
The name XLSTM reminds me of the time in the late eighties when my university professor got accepted to hold a presentation on WOM: write-only memory.
He stated in interviews "Wir werden das blöde GPT einfach wegkicken" (roughly: We will simply kick silly GPT off the pitch) and he just founded a company to secure funding. Interesting times.
Someone gathered most of the available information here: https://github.com/AI-Guru/xlstm-resources
Being a researcher at a public university in a country that doesn't exactly splurge on this kind of research he has to get creative to get any meaningful amount of funding.
So I guess what's left is doing these grand proclamations that you are going to "knock the crown off OpenAI" etc. Though, some sort of vision is good to have for sure :)
The benchmarking done in the table 1 is extremely questionable. Their table basically contradicts the results from multiple peer reviewed papers, especially for RNNs which report results much closer to baseline transformers (and conducted much larger experiments btw).
Page 40 they mention that all models are trained with the same lr for comparability.
> Contradicts their own scaling laws table which uses different lr for different models
> And no it is not a fair comparison to use the same lr to test all these different models. Benchmarking results just looks like they are using tuned hyperparameters for their model which happens to not work for other models.
RWKV-v6 > RWKV-v5 > RWKV-v4, not the other way round obviously. HGRN 8 ppl worse than baseline transformers? NIPS 2023 spotlight paper btw.
However they completely messed up benchmarking experiments for various RNN models which in their papers claim comparable and even better performance than base transformer.