Alas, it doesn't appear to work well for longer contexts:
https://twitter.com/arankomatsuzaki/status/16390003799784038...
Has anyone here experimented with this recently to confirm?
https://twitter.com/arankomatsuzaki/status/16390003799784038...
Has anyone here experimented with this recently to confirm?
We believe this can be done for 16k to way beyond 100k
Research in how RWKV handle the hidden state shows that it is barely used (imo: <5%??) meaning lots of headroom for scaling context size
(This is actively being experimented on - we dun really know the limit yet)
I'm going to take a closer look :-)
That sounds promising. Maybe scale (of model and training samples) is all you need.
And RNNs are obviously so much more efficient at inference.
I'm going to take a closer look :-)