I'm a regular involved with the RWKV community.
AMA, on RWKV, and I will do my best to answer them here for the next hour (one at a time)
PS: you can find our discord here : https://discord.gg/qt9egFA7ve
I'm a regular involved with the RWKV community.
AMA, on RWKV, and I will do my best to answer them here for the next hour (one at a time)
PS: you can find our discord here : https://discord.gg/qt9egFA7ve
One of the biggest issue for testing all of this, is it takes a crap ton of GPUs to prove all the alternatives to transformers beyond 1B param.
For example I’m waiting for someone to do a 1B-14B text based diffusion network
Finally, if this is truely the case (and all that really matter is size+dataset)
We really should use an architecture that is cheaper to train and run. And that’s what RWKV represents here
You can even run the 7B quantized model reasonably on most laptops (try the rwkv-cpp / rwkv-cpp-node project)
- it is NOT backed directly or owned by any VC funded company
- it is 100% OSS driven by the community (Apache 2 license)
- it’s currently the top OSS chat model that can be used commercially on the chatbot arena score board
- IMO it is undertrained, so expanding the training data alone will make it much better (however for the sake of this paper, we wanted to focus on architecture not training data, so we compared similarly trained models)
And yes we do have multiple experiments and plans to make it better. It’s a list, and we will not know which is final until we try. Individual members can go to great lengths on what they are working on
For better or worse, being truly OSS means our initiatives are more disorganized then a centrally planned org
- integrating this with AI platform X/Y/Z
- setting up evals
- improving the code quality
- making a how to guide (it’s stuck on my todo list)
- helping with dataset
- doing silly experiments on how the architecture work (and if the changes give good result)
- etc etc
One of the community goals is to make this a model for EVERYONE on earth that means we need quality dataset for all the non English languages
So even on that level there are things to do
( find something that interest you on the community )
To be fair, that filters the majority of models in the scoreboard.
Where we have more OSS models to choose from without weird rule lawyering gotchas. Or needing to be from a research institute / a license to download the weights
For example, I'm currently attempting to use RWKV for named entity extraction. I ask it to analyze a piece of text and provide output in JSON format. It starts off great. However, eventually, it seems like the beginning of the JSON list 'overtakes' the question I asked, and it starts to just produce random data that would seem plausible based on the set of things in the list. I realize this is due perhaps to the precision losses of the RNN as weights decay.
However, I feel there ought to be some way we can prevent that. Any thoughts?
Ask the question / explain the task first. Then give it the data you want to extract from.
Also you may want to give a one shot example for best result (the instruct training of raven model is very limited)
I'll say "get a list of Blah from the following document in Json format like this:
Example"
Then I feed the document and add a spot for the answer.
The model begins correctly. But usually in the middle of the Json list generation, it will veer off, and start hallucinating as if it forgot the document and the task. I'm happy to share specifics and datasets but this is a cross cutting problem.
Rwkv is able to answer my questions when I ask simple yes/no or classification. It's the listing that throws it for a loop. Transformers do not have the same problem. Both llama and gpt are able to maintain focus.
Also, do you know where I'd find information on how the current weights were trained?
(You are using raven right? That’s the instruct trained varient)
Btw ping the discord if ur looking into finetuning for your usecase
Unfortunately I really would like machine readable responses and raven is a bit too verbose.
Looking at fine-tuning right now.
Most of the focus is in the 1-14B range. Due to constraints of the dataset sizes (chinchilla law), and GPUs available
Community demand is also mostly in this range as there is a strong desire to optimise and run on local GPU. So more focus is in this range.
Not representing blink directly here - but if anyone wants to see a 30B / 65B model. Reach out to contribute the GPUs required to make it happen
The code is already there, just need someone to run it,
Ps: I too am personally interested in how it will perform at ~60B, which I believe will be to be optimal model size for higher level of thoughts (this number is based on intuition not research)
You might find that thread interesting, they're taking submissions for potential partnership with LambdaLabs a cloud compute company that has a few hundred H100s laying around. They have an open form and their cofounder is currently doing the rounds having meetings and this may be a good candidate.
I'm not associated with them at all, just interested in the space and things going on.
what datasets are available to the community, how big are these, are they needed to be updated from time to time, where are these stored, what are the usual cost ranges involved, ...? :o
thank you!
If not, you are getting diminishing benefits for each param you add
I’m extreme cases your model can even perform worse with more param due to lack of training data
More complicated: the quality of the data matters as well
So there are 2 major directions. Build efficient models with good dataset and optimal param count for the task
Or go big on everything (aka openAI) which requires monster GPU time for every reply token
There are obviously in between as well. Hence why the question is so loaded
Ballpark: if your not setting aside a 100k for GPUs alone, to train a 60B model from scratch, your probably not ready to train one
8 x 8 x 8 A100, should be able to do a 100k++ tokens/s at that size
With a dataset of 1.2 trillion tokens. That’s 12 million seconds. Or 140 days
(PS: this is why everyone is training <60B, its crazy the cost, even if my math estimate is wrong by 300%, its still a crazy number)
Is there some kind of dedicated fund for training hardware? Donating an A100 sounds unlikely, but surely they could be crowdfunded?
If you want to help fund RWKV, the ko-fi link is - https://ko-fi.com/rwkv_lm
IMO: this needs way more funding, just to sustain blink leading this project, let alone GPUs for training.
(Also - current tests shows this model doing really badly with 4bit quantized, but alright at Q5 and Q8)
We are talking about over 10x reduction in GPU time for inferencing tokens and for training too
Aka it’s cheaper and faster
Alignment is frankly IMO purely a dataset design and training issue. And has nothing to do with the model