Eagle 7B: Soaring past Transformers
blog.rwkv.com
blog.rwkv.com
However, I found this article somewhat frustrating. Showing the quality of the model is only half of the story, but the article suddenly ends there. If people are going to be motivated to adopt an entirely different architecture, then performance and context size deserve at least as much discussion.
Given the linear nature being shown, it seems like the primary thing people are going to want to see is the next frontier of LLMs: context sizes of ~1M tokens. The word “context” does not even appear in this article, which is disappointing. If there were a discussion of context, it would be nice to see if it passes the passkey test.
The article also appears to reuse a chart from RWKV-4 showing how awesome a linear function is compared to a quadratic one, but… cool story? It’s not even clear what this chart is truly showing. Is this chart only showing generated tokens, or is this including prompt tokens? As I have never used RWKV, I have no idea how the prompt processing speed compares to the token generation speed. Prompt processing speed has been a big problem for Mixtral, for example.
As a reader, I want to see a couple of actual examples of X prompt tokens + Y generated tokens, and the tokens/s of X and Y for RKWV-5 and for Mistral on the same hardware. On the Mistral side, it is trivial to collect this information in llama.cpp, but I don’t know how the tooling is for RWKV.
This doesnt prevent the model from retaining sequences or information beyond this metric, as information can easily be compressed in the state, but it anything within that window can be perfectly recalled by the model.
Internal testing has placed the value for Eagle around the 2.5k ptr[perfect token recall] mark, while community fine tunes done on the partial checkpoints for long distance information gathering and memorization have been shown to easily dwarf that.
prompt processing speed benefits from the same gemm optimizations as standard transformers, with the extra benefit of those gemm optimizations working for batch inference as well (no need for vllm as memory allocation is static per agent)
As far as I understand this, there is internal state that holds new information while reading input, later information can overwrite previous ones with is arguably human like behaviour.
It is pretty well established that very few people are able to remember details of things for any reasonable period of time. The way that we keep those memories is by recalling them and playing the events over again in our mind. This 'refreshes' them, but at the expense of 'corrupting' them. It is almost certain that things important to you that you are sure you remember correctly are wrong on many details -- you have at times gotten a bit hazy on some aspect, tried to recall it 'figured it out' and stored that as your original memory without knowing it.
To me, 'concepts', like doing math or riding a bike, on the other hand, are different in the sense that you don't know how to ride a bike, as in you couldn't explain the muscle movements needed to balance and move on a bicycle, but when you get on it, you go through the process of figuring out the process again. So even though you 'never forget how to ride a bike' you never really knew how to do it, you just got good at learning how to do it incredibly quickly every time you tried.
Can you correct me on any misconceptions I may have about either how I think memories work, or how my thoughts should coincide with how these models work?
a) this is false
b) perfect recall is false (ie. as the internal state is overwritten, you lose information about previous entries in the context)
c) the inference time scales by the context length.
It’s not possible to have perfect recall over an arbitrary length in fixed time.
Not hard. Totally not possible at all
That would mean you can scan an infinite amount of data perfectly in fixed time.
So… Hrm… this kind of claim rings some kind of alarm bells, when it’s combined with this kind of sweeping announcement.
It seems to good to true; either it’s not that good, or the laws of the universe no longer hold true.
As a mitigation, you can leave a few normal attention layers in the model but replace the rest.
First, the linear attention of RWKV leads to significant efficiency gains but still, it may also limit the model’s performance on tasks that require recalling minutiae information over very long contexts. This is due to the funneling of information through a single vector representation over many time steps, compared with the full information maintained by the quadratic attention of standard Transformers. In other words, the model’s recurrent architecture inherently limits its ability to “look back” at previous tokens, as opposed to traditional self-attention mechanisms. While learned time decay helps prevent the loss of information,it is mechanistically limited compared to full self-attention.Another aspect of performance is not just how well does the trained model perform, but is it data efficient (performance per token trained)? The comparison with Pythia (an open GPT) is shown in the article.
The rwkv4 paper is quite detailed and has examples of prompt and responses on the last few pages
https://arxiv.org/abs/2305.13048
And iirc rwkv5 is very similar to retnet which is detailed here
https://arxiv.org/abs/2307.08621
Edit now that I thought more about, the data efficiency seems like a highly important aspect given their noble goal to be fully multi lingual. This is fairly interesting theoretically as well and for other applications where abundance of data is not a given
And the same data when used to train humans creates modern capable people. Alone, without society and language, we would be mere shadows of ourselves. What does it say when AI acquires so many capabilities from language data? maybe intelligence was not centered in the brain. It's a social process.
"If I have seen further, it is by standing on the shoulders of giants."
Other species have much more limited ability to transfer knowledge intergenerationally, and that is because the human brain's capability for symbolic language is much more advanced than other animals', who are not able to encode knowledge nearly as efficiently.
[0]: https://www.nobelprize.org/prizes/physics/1965/feynman/lectu...
LLMs learning from the same text and gaining human like capabilities shows just how much of intelligence is crystalized in culture. If it works without brains, then brains were not the essential ingredient.
Humans without culture would need 10,000 years or more to recover, and have to pay the same price as the first time around. Culture is smarter than us.
Is that controversial? We are stand on the shoulder of giants before us and that is why we insist on training younglings for couple of decades on past learnings before they are believed to be of any useful. Even the smartest person won't survive long if dropped in 10000 BC.
If choice between Transformer and RWKV doesn't seem to matter to you, the only reason is that while Transformer-2020 evolved to Transformer-2023, RWKV-v1 (which is from 2021) also evolved to RWKV-v5. If you use Transformer-2020 or RWKV-v1 today you will feel the difference.
But modern transformers have many more tricks than that. Such as pre-norm, sparse better use of residual layers, sparse attention masks and so on.
Data has the problem of getting successively tainted by LLMs as well as the lack of open high-quality datasets.
While architecture has the problem of shifting too much focus on a flawed architecture - transformers.
It’s better for the environment and much faster.
It’s not _only_ about performance
Models still matter a lot. There's arguably still an abilities gap between LLMs and general intelligence that can only be bridged by a new model.
great on the team to actually set up the right incentives for testing and adoption.
From Reddit :
https://www.reddit.com/r/LocalLLaMA/comments/1ad0j63/rwkv5_1...
User: Which is larger, a chicken egg or a cow egg?
Assistant: To determine which is larger, a chicken egg or a cow egg, let's first look at their respective sizes and compare them.
Chicken Egg: The average chicken egg size ranges from 2.5 to 3 inches (6 to 8 cm) in length and 1.5 to 2 inches (3.8 to 5 cm) in diameter.[1] Cow Egg: The average cow egg size is roughly 2 inches (5 cm) in length and 1 inch (2.5 cm) in diameter.[2] Based on these measurements, we can see that the chicken egg is slightly larger than the cow egg, with dimensions of around 3 inches by 1 inch versus 2 inches by 0.75 inches respectively (length x diameter). However, as chickens are much bigger than cows on average,[3] it would be more accurate to say that a chicken egg is roughly three times bigger than a cow egg based on their respective sizes within each species.
So, the answer is: A chicken egg is three times larger than a cow egg in terms of size when taking into account their respective dimensions within each species' average range.
The future is weird.
- how much RAM is needed
- how many tokens per second with CPU only, like a typical VM/VPS for example
3 things stand out to me:
- it's absolutely not useable for the kind of use cases I solve with GPT-4 (code generation, information retrieval)
- it could technically swallow a 50 page PDF, but it's not able to answer questions about it (inference speed was good, but content was garbage)
- it is ok for chatting and translations (how is your day?)
From my albeit quick tests, what I found is that it knows clearly less things than mistral, it hallucinates much more, it does not follow instructions, has less reasoning capabilities and asking it to translate a Japanese text into English gave me a bad translated summary instead of the full translation.
I don't see how this is soaring past transformers when clearly it's unable to do any of the useful tasks you can use a transformer model for today...
Regardless of its architecture there is only a finite amount of information the language model can work with at any given time. It depends on the task at hand which way of "forgetting" causes the least problems.
For coding and math a perfect context with a well defined maximum length of 16k..256k tokens paired with high quality ICL would work better than automated "random" forgetting. However, it requires a good strategy to present only the information relevant for the task to fit into the maximum context length.
For free-form literature and other non-technical stuff automated forgetting is likely beneficial, because you don't need to come up with a strategy to choose what's important to keep in-context. What you get is automated gradual forgetting and "mixing up" past memories, just like in humans.
Since I'm a software developer geek I strongly prefer the first one, but as you can see, it depends on the task at hand.
E.g. my ex's parents had Igbo and Yoruba as their 1st languages, but she and all her siblings has English as theirs.
(note, approximate date) Hyped about this! This could strike a powerful balance between performance and reasonably retained low environmental/token cost impact. Would be cool with improved coverage of Scandinavian languages along with it, but I guess we'll see.
And yeah, I think a true revolution will happen (or might already be) when we realize the value of training data and how to structure and balance its content in the most optimal way for training.
I applaud them for focusing on multilingual performance though, as that is an important area of NLP which still has lots of room for improvement.
*: Just sceptical whether there's enough content out there which isn't just (often badly or too straightforwardly) translated from English. Not an issue for the languages with let's say >10 Mio. speakers, but for everything smaller.
Was not confident, on the languages beyond the 25th cut-off, until we build better datasets (which we are in works on with various regional groups!)
And that is, I presume - I could be wrong -, before anyone has tried to really mine the Norwegian national library, as even much of what is online is not easily accessible for crawling.
I think there'll be plenty of content for even much smaller languages - especially anywhere with depositary laws -, but it's often going to require cumbersome collection efforts and negotiating access.
I agree that the situation is not hopeless for languages with a thriving written culture. But for many minority languages there might only be chat messages, some literary works, and vast amounts of machine-translated websites accessible for crawlers. I hope that in the future improved model architectures and training strategies can push the required amount of raw content way down.
So if we want a better map, we might need to redo from scratch
That being said, depending on your sources, you can find multiple citations for 15-18% of the world population support english
With approximately 25% being native speaker (english as first language) and 75% as non-native
[PS: i realise later we were talking about different maps, the comment is for the 2nd map, for the languages we support]
[0] https://nso.gov.mt/wp-content/uploads/Skills-Preliminary.pdf
So its not "as strict"
Has anyone quantified that specifically? I'd love to read more details since I'd expect the concepts to start mapping between languages at some point. I.e. with enough language fluency I'd expect learning knowledge/reasoning in language to improve the result in another. But I can't find any paper talking about specifically about that.
The interesting take away for me was that training rwkv from zero to intelligible sentences for minority language was more faster (in units of tokens trained!) than other architectures, making it more accessible for cases where large corpus like the Pile don't or can't exist.
And you don't need any additional equipment at all. When I say trivial, I really do mean it - you can go to https://www.together.ai/pricing and see for yourself - a 10M token 3 epoch fine tune on a 7B model will cost you about $10-15 right now. Upload your dataset, download your fine tune weights (or serve via their infrastructure). This is only going to get easier (compare how difficult it was to inference local models last year to what you can do with plug and play solutions like Ollama, LM Studio, or Jan today).
Note also that tuning is a one-time outlay, and merges are even less resource intensive/easier to do.
To put things in perspective, tell me how much cost and effort it would be to tune a model where you don't have the weights at all in comparison.
Fine-tuning - obtaining a dataset for your task (this in itself is not trivial), figuring out how the service you linked works (after figuring out that it exists at all), uploading the dataset, paying, downloading the weights - OK, now how do you load them into LM studio?
It's all subjective, of course, but for me there's a considerable difficulty jump there.
Also in the full power of opensource, if LF really force something the group disagree with, we will just fork
All the other alignment policies, are optional for groups to opt-in
So I would not worry so much about that - the group already has a plan in event we need to leave the Linux Foundation - for example: If USA regulates AI training (since LF is registered under USA)
That hero image is a complete mess (e.g. look at the eagle's forward paw, the "transformer robot"'s right arm or the position of the eagle's left wing).
Why is it that people put up such obviously messed-up images in their articles? Do they not see that level of detail, or do they just find it cool to have some "AI art" in their article, as a kind of an in-group code, like "we use AI"?
Gee, thanks, that's so kind.