The RWKV language model: An RNN with the advantages of a transformer
johanwind.github.io
johanwind.github.io
Its also not really an RNN. The best way to describe the key time mixing operation is a normalized exponentially weighted moving average (EMA)- no non-linearity. Once viewed this way, its not surprising that it struggles at longer contexts- everything decays, and it has limited space to put things. Of course, it does have some clever tricks, and can choose to remember things for a while by upweighting them, but not forever.
The current memory techniques we have outside of ultra long context lengths are lossy and imperfect. I wish the langchain spammers (yes it's a good tool) would acknowledge this when they keep posting everywhere about the "memory module".
I wonder if RLHF would boost performance for this architecture in the same way
They have a long lookback window (30,000 tokens in GPT4!) so the issues are smoothed over.
LSTMs end up being very finnicky to train and their sequential nature makes model scaling bottlenecked in parallelism.
I'm excited about the next iteration of models focusing on inference. RWVK seems promising on this end. Also love the citizen science aspect of RWVK coming out.
AGI fears Fabrice Bellard
This seems to be made for mass parallelism.
Question: Would this scale to 100 million of collaborating 4G/5G smartphones? We have operational federated learning code and decentralised learning code (trivial stochastic gradient descent). Application: simple learn-to-Rank based on ClickLogs like its 2005 [1]. Then you have a true decentralised Google.
Would love to link RWKV to other pure decentralised tech. In the past we have build the first self-compiling Android app and first Android-to-Android P2P overlay network. Without any helper peers for carrier-grade NAT puncturing. My university systems lab lacks the size to keep up with the recent pace of innovation.
He also wrote an engine: "LibNC: C Library for Tensor Manipulation"
> [ ...]/LLMs
That could be a wish come true.
Edit: in fact, already one of his marks can be seen: «Larger models work optimally on lower cost GPUs (e.g. RTX 3090, RTX A6000) thanks to efficient quantization»
Prompt:
A website can be built in 10 simple steps
Output:
1. Research
2. Research
3. Research
4. Research
5. Research
6. Research
7. Design
8. Design
9. Design
10. Design
Does anybody have some examples of output that it generates to explain what this is capable of?
I tried a few queries. It works okay for basic prompts for coding questions but can give pretty good results with more context / prompt engineering
Is there a reason the creator hasn't written or tried to publish a paper on it? I'd love to see an peer-reviewed discussion detailing how it works, maybe studying how it behaves internally, evaluating its performance more rigorously than posting a few examples, maybe in comparison to Transformers with similar parameter counts, or training/inference costs in FLOPs. Perhaps that's not the creator's main priority, though.
I kinda respect that, can't force someone to explain their ways if they prefer hacking to writing. I would also love to see more rigorous measurements and discussion though.
The creator replied: "Thank you :) Too busy for that at this moment, but I will get a paper out later this year." once on a reddit post asking exactly this.
https://old.reddit.com/r/MachineLearning/comments/1135aew/r_...
What's the benefit of RWKV then? The results are practically identical to transformers.
- only need the hidden state at position t to compute t+1, i.e. run inference locally on edge devices
- fine-tunable to longer context lengths than seen in pre-training
If it continues to scale as well as it has so far it's pretty huge.
> So we basically have a model which trains like a transformer, except that long context length is not expensive. And during inference, we need substantially less memory and can implicitly handle “infinite” context length (though in practice, the model might have a hard time generalizing to much longer context lengths than it saw during training).
The tl:dr Linear Attention avoids the quadratic token scaling and is equivalent to specific instances of RNNs.
30B/q4 requires 20GB of RAM while 3090 has 24GB.
Llama 30B 4-bit has amazing performance, comparable to GPT-3 quality for my search and novel generating use-cases, and fits on a single 3090. In tandem with 3rd party applications such as Llama Index and the Alpaca LoRa, GPT-3 (and potentially GPT-4) has already been democratized in my eyes.
Need 64GB for 30B
And 128GB for the 65B
I'm not sure about the performance, but I think it should be ok? Especially given how much Apple has been investing in what -- I believe they call -- neural cores?
Here's some more context: https://news.ycombinator.com/item?id=35105364
What benefit does that offer over 2 GPUs?
Meanwhile the A6000 VRAM bandwidth is 768 GB/s (= 16 Gb/s (GDDR6) × 384 bit-width ÷ 8 bits per byte).
(Also, the RTX 3090 has faster VRAM, >900 GB/s, than the A6000, because it is GDDR6X.)
I hope this only man on the planet who can train big RNN models will beat OpenAI one day, like he said he would.