A lot of information is wrong, a lot is only true in context (e.g. 2+2=5 features prominent in the book 1984), even more is spam or machine generated.
If you see the garbage that goes in, I am always amazed at how these models do what they do.
A lot of information is wrong, a lot is only true in context (e.g. 2+2=5 features prominent in the book 1984), even more is spam or machine generated.
If you see the garbage that goes in, I am always amazed at how these models do what they do.
The sqrt(-1) sometimes doesn't exist, sometimes it's 1i. 2+2=4, except in literature where it can be 5. 1+1=2, but sometimes 3 in advertisements or in ironical text.
We often have some ideas about e.g how it works in a quiz, where you know there is only one factually correct answer. And we are disappointed if the model is wrong. But even in a quiz setting the jury gets that balance wrong every so often, where there are other answers than the official one which are also correct.
Even "logically valid" is context dependend. This is not to say that models don't hallucinate, just that even within the logically valid answers, there is hidden context surrounding the data which is not expressed in the data itself. Fermats last problem is a solved problem in mathematics, but not in documents from before 1994.
1984 having 2 + 2 = 5 makes sense in context as a human reading the book, and ChatGPT dot-producting the book can also compute the context and not say that 2+2=5.
ChatGPT's not Mathematica, and we already have calculators. My hammer is terrible for driving in nails, so I don't use it for that.
(Idea assessment in general. Handled ideas in thought processes are still input.)