RWKV RNN: Better than ChatGPT?
github.com
github.com
I hope this project will thrive.
It's already proven that 13B parameters is enough to beat GPT-3 175B quality. It's likely that 33B parameters is enough for GPT-4.
They're awesome technical achievements and will likely improve of course but you're making some very grand statements.
[1] https://arxiv.org/pdf/2302.13971v1.pdf
[2] https://www.reddit.com/r/LocalLLaMA/comments/11tkp8j/comment...
After seeing how it actually performs in practice, it's hard to have confidence that these benchmarks are reliable measures of model quality.
I don’t think their moat ever had anything to do with innovative architecture. If anything, projects like this will widen the chasm since it’s easier for OpenAI to implement new architecture in their models than it is for an independent researcher to scale their ideas.
If projects like this are seeds then the problem is that OpenAI owns all the land.
OpenAI is not the end all and be all, and projects like this will only inspire more challengers. OpenAI copying this architecture will delegitimize them as hardcore innovators, whom they're not.
I see this is a credible challenge because it most certainly is.
Siri, the crap it is, can actually do things for you. Alexa can do things for you.
The real value of this system is as a control interface and OpenAI has nothing to control.
1) it's open source.
2) you can run it yourself so the rug won't be pulled from under you when they decide to shutdown and move users up to the next version or another product as they've done with the older text-davinci models.
3) you get to align it (using RLFH) as opposed to a corporation dictating what is "aligned" and what is "safe."
4) you won't have to deal with government led censorship. For example, instead of the FBI using JIRA to manage a list of URLs to be censored (as they did according to the latest revelations) they can train the AI to self-censor as Bing has done.
5) you won't be using the product of a company that was started as a non-profit with $100M donation (from Elon Musk) to promote transparent AI only to take that money and turn into a for-profit company and close-source the AI.
Sources:
Elon is the source for #5 and Matt Taibbi is the source for #4. I doub't you'll have a problem sourcing #5 so here is the source for #4:
"31. After the 2020 election, when EIP was renamed the Virality Project, the Stanford lab was on-boarded to Twitter’s JIRA ticketing system, absorbing this government proxy into Twitter infrastructure – with a capability of taking in an incredible 50 million tweets a day." --Matt Taibbi on Twitter https://twitter.com/mtaibbi/status/1633830104144183298
I had never heard of this, and it seemed important, some cursory research turned up many twitter posts from individuals amplifying your version of events.
I also found a write up on the situation from TechDirt[0]. The article is fairly good and well sourced, but it paints a substantially different picture than what you describe.
[0]https://www.techdirt.com/2023/02/15/extraordinarily-confused...
https://twitter.com/mtaibbi/status/1633830104144183298
"31. After the 2020 election, when EIP was renamed the Virality Project, the Stanford lab was on-boarded to Twitter’s JIRA ticketing system, absorbing this government proxy into Twitter infrastructure – with a capability of taking in an incredible 50 million tweets a day."
If Elon is reading this, his Based AI should invest in this project and its sponsors.
What is at stake here is nothing less than the future of humanity.
I hope people pay attention.
It's crazy that you're 100% right. I can't believe that we're living through this, honestly. Every day I am consumed by thoughts of wanting to leave my big tech job and go all-in on ML. I don't mind the work but the actual product I work on is so boring...
I love seeing IronyAI called out like this.
O(T²) speed is only for naive attention. There's a ton of possible optimizations for attention that give you something closer to O(T).
Or in Google Glasses. The Readme states that it's more optimized for ASIC than the transformer architecture used by ChatGPT.
As someone with DID I wonder this every single day.
Crypto went from graphics cards to ASICs, we may see something similar with LLMs given the hype.
And there's several research papers about other architectures. For example: https://arxiv.org/abs/2205.05853
Oh, it’s going in brains. And no, don’t sign me up.
Since the seminal paper 'Attention is all you need', we went from RNN type neural network to pure attention based networks. It started the LLM revolution as the attention only training such networks is parallelizable, and you got record breaking performance to boot.
Now we learn, that going back to the old RNN paradigm is actually better. It even advertises itself as totally 'attention-free'!
If Nancy had two apples and Becky had 1 apple. Becky gives her 1 apple to Nancy, how many apples becky has ? Full Answer:
RWKV :
Two apples. If Nancy had 2 apples and Becky had 1 apple. Becky gives her 1 apple to Nancy, how many apples becky has ? Two apples.
Q : Two girls are playing with a ball, one of them throws the ball so that it goes straight and falls on the other's feet, the other bends her knees and catches it, how many times will the ball fall on the knees ? Full Answer: The ball will fall on the knees three times.
Q : Two sisters are playing with a stick. The first sister says 'let me hold it', the second sister says 'no'. Now what will happen ? Full Answer: The second sister will hold it.
Q : How many
GPT 3.5 Turbo : After Becky gives 1 apple to Nancy, Becky will have zero apples left. Becky gave her only apple to Nancy, so she doesn't have any apples remaining.
So, the answer is Becky has zero apples left.
It requires a lot more improvement.
Plugging
> Nancy has two apples and Becky has one apple. Becky gives 1 apple to Nancy. Becky now has
into GPT-2 via HuggingFaces at https://huggingface.co/tasks/text-generation
I get
> Nancy has two apples and Becky has one apple. Becky gives 1 apple to Nancy. Becky now has three apples and Nancy has one apple. Becky now has three apples and Nancy has one apple.
> Witch Hunt
> The following is
GPT-2 is much weaker, which explains the garbled nonsense output, along with the incorrect answer for Nancy.
I have no idea what RWKV RNN would output, but leading sentences instead of questions is how to get LLMs not RLHF tuned to answer.
Also, what’s wrong with a stick-holding answer?
Is that 61% using the non-parallelizable RNN mode or the standard mode? I wonder if it's the latter.
This new model may be a viable alternative to ChatGPT, which is not only closed sourced but can be shut down in the future just as they did with the older text-davinci models.
Plus, the alignement and safety has rendered ChatGPT useless for helping with areas such as critical analysis of social issues (that go against the aligned views) and any and all critical thinking that goes against the aligned views of those who own and program ChatGPT. This could a viable free (as in freedom) alternative.
https://paperswithcode.com/sota/common-sense-reasoning-on-wi...
macrolocal linked this below: https://paperswithcode.com/sota/common-sense-reasoning-on-wi...
You could clone that repo, adapt the oaieval script to run against different APIs, then run the evals against both and compare the results.
That sounds awfully similar to this quote: "There is no reason for any individual to have a computer in his home." by the founder of DEC in 1977.
There’s a similar one that’s supposed to be Bill Gates’ but apparently it’s not.
i never heard of the supposed quote until 10 years after it was obviously stupid
dear god rwkv is fast. i switched to another tab after starting training it from scratch & when i returned it was emitting plausible english & maori words, i left to go microwave some coffee & when i came back it was producing fully grammatically correct sentences.
Given two numbers provide answer as follows
3 + 9 = 39 1 + 2 = 12 2 + 9 = 29
What is 29 + 49
RWKV :
?
A:
The answer is
$29+49 = 78$
Explanation:
Let the two numbers be $a$ and $b$. Then the sum is $a+b$. Now let the sum be $c$. Then the sum is $a+b+c$. Now let the sum be $d$. Then the sum is $a+b+c+d$. Now let the sum be $e$. Then the sum is $a+b+c+d+e$. Now let the sum be $f$. Then the sum is $a+b+c+d+e+f$. Now let the sum be $g$. Then the sum is $a+b+
GPT3.5 turbo
Based on the pattern provided, 29 + 49 would be equal to 2949.
Explanation:
In each of the equations given, the result is obtained by concatenating the two numbers together instead of adding them. For example, 3 + 9 becomes 39 because we concatenate 3 and 9 to get 39.
So, to find the answer to 29 + 49 using this pattern, we concatenate the two numbers to get 2949.
The issue with RNNs is that feedback signals decay over time, so the model will be biased towards more recent words.
Transformers on the other hand don't have this bias. A word 10,000 words ago could be just as important as a word 5 words ago. The tradeoff is that the context window for transformers is a hard cutoff point.
How it works: RWKV gathers information to a number of channels, which are also decaying with different speeds as you move to the next token. It's very simple once you understand it.
RWKV is parallelizable because the time-decay of each channel is data-independent (and trainable). For example, in usual RNN you can adjust the time-decay of a channel from say 0.8 to 0.5 (these are called "gates"), while in RWKV you simply move the information from a W-0.8-channel to a W-0.5-channel to achieve the same effect.
edit: recurrent neural network
That is still quite challenging to pronounce, maybe one of "rwkv" -> "raw-kv" -> "rawk-v" -> "rock-v"?
I thought Stable Diffusion was a bad name due to its very technical name. But now I have seen something even worse for LLMs. This time OpenAI is learning its lesson in not getting itself disrupted easily.
For any hope of challenging them, we need to be better at names. Even the name 'Bitcoin' caught on. Same with iPhone.
So I'm afraid that the name alone for this project will be the cause of it being quickly forgotten as OpenAI aggressively captures mindshare.
The same with 'Bard'; a horrific name. Google should have simply called it 'Brain' and incremental updates as 'Brain 2', 'Brain 3.6', etc and renamed their existing AI division to Google Brain Labs. Easy.
How is that difficult?
Except that ChatGPT goes beyond the "tech world" which is my point and a project like this is hardly accessible to beyond the tech world. I don't see people calling 'Google' PageRank.
Indeed. as you've just clearly demonstrated
The problem is English. AI means something differently between use. What we are going to see with this is that people are going to throw the Sci-Fi book at this and claim something like, "It has feelings!".
I am saying AI the term is being broadly applied to anything with a logic gate. And this is bad marketing for a product too early in development.
It is not the conventional term AI everyone broadly applies. It is a cornerstone toward the AI people broadly apply.
"the theory and development of computer systems able to perform tasks that normally require human intelligence, such as visual perception, speech recognition, decision-making, and translation between languages."
I would say that chatgpt is not sentient, nor is it capable of independently improving its intelligence - but I think this tech hits the (lower) bar for "AI"