Yeah that's pretty much what gwern argues here[0]. Or to adapt another proverb: to predict the next token you first need to model the universe.
[0] https://gwern.net/scaling-hypothesis#gwern-difference--effic...
[0] https://gwern.net/scaling-hypothesis#gwern-difference--effic...
Exactly. The "most likely next" series of tokens, for example, when given the first half of a correct mathematical proof, is the correct rest of the proof. I have never seen anyone define "most likely next token" in such a way that this isn't true.
Either humans are not capable of intelligence or computers are capable of becoming intelligent. Neither or both.
that feels like a misunderstanding of how the loss function behaves when used within a sequence