Just one example - it's very common to see ChatGPT and the like respond with "you're absolutely correct! Great insight" to something that is a complete misunderstanding.
Just one example - it's very common to see ChatGPT and the like respond with "you're absolutely correct! Great insight" to something that is a complete misunderstanding.
> It’s been proven that when a model is trained on large volumes of highly factual and non-theoretical data, it learns to always have an answer. DeepSeek V4 Pro (1.6T params, 49B active, 44 AA Intelligence Index score) has a ludicrous 94% hallucination score on the AA-Omniscience benchmark, meaning on questions that it couldn’t figure out, it only stated that it didn’t know around 6% of the time, and the rest it confidently hallucinated an answer. GLM-5.2 scored a 28% hallucination rate, Opus 4.8 was 36%, Fable 5 was 48%, and GPT-5.5 was 86%.
https://arrowtsx.dev/bigger-models/
I think even a 5% hallucination rate would be terrible for a teacher, who should generally be comfortable with saying "I don't know off the top of my head but here is how to find resources to answer your question".
---
So, just to drive the point home, Codex has an 86.9% hallucination rate on the AA-omniscience score in this index https://benchlm.ai/models/gpt-5-3-codex - if you ask it something that wasn't sufficiently covered in its training data, it will confidently make up an answer nearly 87% of the time.
While you might think it is happy to correct you when you are wrong, you don't know that for sure since you don't know when you're wrong. Codex may have been happily agreeing with you about things you had completely backwards.
Yes I am certain that it feels that way. However empirical testing holds a lot more weight than anecdotes.
> The answer for hallucinations is fairly simple: give it facts and tools to ground itself.
The entire danger here is that it hallucinates when you don’t know the ground facts. After all, you don’t know what you don’t know.
So you make it demonstrate the ground facts. Show formal proofs, link to and reference specific documents and fragments in a document database. It wouldn't really be that hard to build a harness with specific support for these things (e.g. performing a tool call that then has UX to highlight fact-checked claims and link to details/references).
That's a great way to get you to listen because your guard is down. Imagine if it told you you were an idiot and then corrected you.
I certainly have, too, but there is still a difference between a person who has a factually incorrect but consistent worldview and an LLM which simply reflects the worldview of the user or even changes between queries.
I don't think creationists have any business being in schools either, for what it's worth, but I think it's easier for a teenager to sort out "Mr. Smith has no clue what he's talking about" vs "I have no clue what's true because the LLM everyone expects me to learn from just confirms everything I ask regardless of what I'm asking".
> Most students lack intrinsic motivation
Woah, this a wild claim. Plus it is on a spectrum per subject. Do you have any evidence to support your bold claim?