Suggesting that an awful lot of calculations are unnecessary in LLMs!
Suggesting that an awful lot of calculations are unnecessary in LLMs!
The obvious way to deal with this would be to send forward some of the internal activations as well as the generated words in the autoregressive chain. That would basically turn the thing into a recurrent network though. And those are more difficult to train and have a host of issues. Maybe there will be a better way.
Hi! I lead interpretability research at Anthropic.
That's a great intuition, and in fact the transformer architecture actually does exactly what you suggest! Activations from earlier time steps are sent forward to later time steps via attention. (This is another thing that's lost in the "models just predict the next word" framing.)
This actually has interesting practical implications -- for example, in some sense, it's the deep reason costs can sometimes be reduced via "prompt caching".
If you want to be precise, there are “autoregressive transformers” and “bidirectional transformers”. Bidirectional is a lot more common in vision. In language models, you do see bidirectional models like Bert, but autoregressive is dominant.