Qwerky: Attention is not what you need? RWKV mashed into QwQ models
substack.recursal.ai
substack.recursal.ai
Their method:
* Keep the old model as a teacher * Copy it, freeze all layers * Rip out the transformer layers * Inject new RWKV layers there * Train against og model as a teacher * Unfreeze entire new model, train against teacher * Train for long context
Their tests were done with a tiny number of tokens for training, like 200-500m total. Results: interesting! It’s definitely better than QwQ on some tasks, and worse on a few.
The super interesting result here is that it seems as if most of the knowledge and intelligence is kept in the feed forward part of the network, not the transformer, since it can be replaced, and the replacement can be trained relatively quickly. That’s super super fascinating.