Andrej Karpathy on Hallucinations
twitter.com
twitter.com
We start with quasi-random initialization of network weights. The weights change under training based on actual data (assumed truthful), but some level of randomization is bound to remain in the network. There would be some weights which settle with low error margins, while some which would have wide error margins. Wider error margins are indicative or higher initial randomness remaining and lesser impact of the training data.
Now when a query comes for which the network has not seen enough baking from the training data, the network will keep producing tokens based on weights that have a larger margin. And then, as others have noted, once it picks a wrong token, it will likely keep on the erroneous path. We could, in theory, however, maintain metadata around how each weight changed, and use that to foresee how likely would the network confabulate.
[1]: https://transformer-circuits.pub/2022/toy_model/index.html [2]: https://rome.baulab.info/
Early on, people assumed a 1:1 relationship between genes and traits. Some are indeed, but many phenotypes are smeared across thousands of genes in a way that is fully mysterious to us.
Sounds like a cop out.
The key takeaway is that we can't trust LLM output and systems need to be designed accordingly.
Eg Agents will never be able to securely take actions autonomously based on LLM responses.