Yes you are, regarding LLMs at least. Here's why:
just for fun I asked a reasonably smart LLM to ...
[be] capable of understanding if it's gone off on
a hallucinatory path ...
"Smart" in this context is a subjective value judgement.
"Hallucinations" are only experienced by living organisms.You then went on to state:
> If you know something rare and the LLM does not, you'll immediately see when it's hallucinating an answer or answering factually.
Again, "hallucinating" is not something an algorithm can do. Also, determining factuality is again subjective based on the person assessing the information.
https://artificialanalysis.ai/leaderboards/models
You will note that in my original comment I very specifically said "deepseek v4 flash 0731", which by many benchmarks/metrics, is "smarter" (again, this is a metaphor) than a smaller or older model. It is also specifically known to be relatively capable of producing code and formulas on demand in several common languages.
And "hallucinating" to mean "outputs plausible sounding gibberish that doesn't hold together consistently". Of course there's no actual hallucination going on.
In this media (comments in HN threads), all I can do is interpret what people write. ;-)
> And "hallucinating" to mean "outputs plausible sounding gibberish that doesn't hold together consistently". Of course there's no actual hallucination going on.
This may very well be what you know to be true and I have no reason nor desire to assume otherwise. The problem is... Many people use the word "hallucinating" in this context literally and not metaphorically.
Since I do not know you, how am I to tell the difference?
It is always a joy when a person, such as yourself, finds the irony in my moniker.
Thank you for this.
But in general, yes, the LLM cannot know about concepts that are far outside of its training set. Humans are the same, I would argue. If you add a good amount of your own knowledge into its context, or better yet, into finetuning, you might find it surprisingly easy to get it caught up on that material.
[1] Ahmed, A., Cooper, A. F., Koyejo, S., & Liang, P. (2026). Extracting books from production language models. arXiv preprint arXiv:2601.02671. https://arxiv.org/abs/2601.02671.
You'd be surprised how few digits you need to make a problem that is presumably unique in earth history. For a typical sum, the number of pre-existing answers would need to scale with 10^n lines of text where n is the number of digits. This expands out of control REALLY quickly. A quick guesstimate has you somehow reading out of a literal black hole at n=21 digits if your LUT is on paper, or n=26 digits if you're using modern HDD technology. O:-)