xLSTM code release by NX-AI
github.com
github.com
And realistically, do we need another GPT4 evaluation paper?
That's only true IF you're thinking specifically of the GNU AGPL. The AGPL - more broadly construed - has a pre-GNU history that goes back further. That said, the GNU AGPL evolved out of the original AGPL. See:
https://en.wikipedia.org/wiki/GNU_Affero_General_Public_Lice...
In March 2002, Affero, Inc. published the original Affero General Public License (AGPLv1) for use with the Affero project and made the new license available for use by other software-as-a-service developers.
For researchers, in all honestly that is a very, very good reason to go GPL. If someone wants to profit off of it, it's not that they can't use the code commercially, they are just forced to hire you or pay you to dual-license it.
There's no reason why a company whose stock goes up $10B due to your model can't cut you a few million of that.
This code will allow people yo experiment and see if it is a viable architecture at foundation/frontier model scale.
xLSTM: Extended Long Short-Term Memory - https://news.ycombinator.com/item?id=40294650 - May 2024 (73 comments)
- They scale linearly in complexity, which means with a longer context window they will be faster and cheaper than transformers
- It's been mostly academic as far as I know, only just recently being published. I don't think there's been an opportunity to use them at 'real world scale' yet, although tbh I'm a little uncertain what you mean by it.
Model: #Params (M), SlimPajama (15B) ppl ↓
- GPT-3: 356M, 14.26
- Llama: 407M, 14.25
- H3: 420M, 18.23
- Mamba: 423M, 13.70
- Hyena: 435M, 17.59
- RWKV-4: 430M, 15.62
- RWKV-5: 456M, 16.53
- RWKV-6: 442M, 17.40
- RetNet: 431M, 16.23
- HGRN: 411M, 21.83
- GLA: 412M, 19.56
- HGRN2: 411M, 16.77
- xLSTM[1:0]: 409M, 13.43
- xLSTM[7:1]: 408M, 13.48
There are more detailed perplexity and task benchmarks in the paper. Overall, all the architectures perform very similarly on every benchmark, sometimes xLSTM is slightly ahead but not always, and the difference is not really meaningful.
This is great news though, it means we are not losing anything by switching to xLSTM and we get important advantages like the scalable context window.
I'm quite excited about this because we can potentially have the LLM remember what you say and do few-shot persistent learning from user interaction (updating "itself", the state vector). It would be very interesting if LLMs were no longer static. Although I'm sure it will be a challenge to train the model to keep such learnings in its memory long-term.
The paper: https://arxiv.org/abs/2405.04517
Little bit of a nightmare too. Instructions keep piling up for you that you no longer openly can access and remove
We really don't even know how Mamba vs Griffin compare.