We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.
We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.
If we would know that, there would be no need for interpretability research.
This is not what's meant by the statements that we don't know how LLMs work. Explain why LLMs are so good at programming, finding bugs, and developing mathematical proofs. Like, way better than all prior tools specifically designed to be bug finding tools, despite being merely "language models".
I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?
This is like saying we don‘t know how a car works because a car can beat the best human athletes in 100 meter dash.
> maths are internally coherent and entirely theoretical
Nope. This kind of wish-washy thinking is not what we mean by understanding.
Claiming that you understand LLMs is similar to saying that you understand how our biology work because you understand evolution. No - you understand the mechanism behind evolution, but not the complexity it produces.