Q-Transformer: Scalable Reinforcement Learning via Autoregressive Q-Functions
q-transformer.github.io
q-transformer.github.io
Wow. Pretty darn cool! <3 :'))))
Very mobile-device and battery-powered systems friendly. :')))) ;'DDDD
Just how much compute/memory are we saving here?
My understanding is that a 1BN transformer is about 2BN flops/inference, so about 1TFLOP for a 500 sequence of inferences (and also about several GB of memory)
What would be the equivalent RWKV (let ignore the inevitable loss penalty which could be significant..)
It only requires the previous state.
(there's a discord, you should join it with further questions! I unfortunately am not as informed as I should be on this one, other than the fact that it is _very_ mobile friendly). The performance diff is slight but not too bad really, all things considered. And I think it comes out on top for raw efficiency per parameter/flop, IIRC.
An interesting concept, for sure! :'DDDD :'))))
If this technique is good, I'll wait until I can learn about it without joining the Discord.
But "getting" the idea of Q learning for a small state space is fundamental and surprisingly approachable.
I wrote this extensive tutorial for teaching deep reinforcement learning, with a focus on getting intuition from code. you will find RL theory is heavy on math despite needing math for very little other than abstractly representing some machine goal and intuition, of which code serves a native programmer already very well.
i spent years failing to learn machine learning and RL until i just started reading source code. books of integrals i never ended up needing.
dont be turned away by the joking nature of my tutorials. there is a real depth in there
I'd suggest getting a good book or other teaching resource and solve a few Gymnasium[0] environments. Unlike supervised machine learning, you don't need someone else's data, you generate your own data.