>Looking through the document, I can identify several instances where it's written in the first person:
And it went on to show a series of "they/them" statements. I asked it to clarify if "they" is "first person" and it responded
>No, "they" is not first person - it's third person. I made an error in my analysis. First person would be: I, we, me, us, our, my. Second person would be: you, your. Third person would be: he, she, it, they, them, their. Looking back at the document more carefully, it appears to be written entirely in third person.
Even the good models are still failing at real-world use cases which should be right in their wheelhouse.
Could you give an estimate of how many "dumb errors" you've encountered, as opposed to hallucinations? I think many of your readers might read "hallucination" and assume you mean "hallucinations and dumb errors".
As a user, when the model tells me things that are flat out wrong, it doesn't really matter whether it would be categorized as a hallucination or a dumb error. From my perspective, those mean the same thing.
I haven't been keeping a formal count of them, but dumb errors from LLMs remain pretty common. I spot them and either correct them myself or nudge the LLM to do it, if that's feasible. I see that as a regular part of working with these systems.
It's hard to know why it made the error but isn't it caused by inaccurate "world" modeling? ("World" being English language) Is it not making some hallucination about the English language while interpreting the prompt or document?
I'm having a hard time trying to think of a context where "they" would even be first person. I can't find any search results though Google's AI says it can. It provided two links, the first being a Quora result saying people don't do this but framed it as it's not impossible, just unheard of. Second result just talks about singular you. Both of these I'd consider hallucinations too as the answer isn't supported by the links.
I just got pointed to this new paper: https://arxiv.org/abs/2508.01781 - "A comprehensive taxonomy of hallucinations in Large Language Models" - which has a definition in the introduction which matches my mental model:
"This phenomenon describes the generation of content that, while often plausible and coherent, is factually incorrect, inconsistent, or entirely fabricated."
The paper then follows up with a formal definition;
"inconsistency between a computable LLM, denoted as h, and a computable ground truth function, f"
| AI hallucinations are incorrect or misleading results that AI models generate.
It goes on further to give examples and I think this is clearly a false positive result. > this new paper
I think the error would have no problem fitting under "Contextual inconsistencies" (4.2), "Instruction inconsistencies/deviation" (4.3), or "Logical inconsistencies" (4.4). I think it supports a pretty broad definition. I think it also fits under other categories defined in section 4. > then follows up with a formal definition
Is this not a computable ground truth? | an LLM h is considered to be ”hallucinating” with respect to a ground truth function f if, across all training stages i (meaning, after being trained on any finite number of samples), there exists at least one input string s for which the LLM’s output h[i](s) does not match the correct output f (s)[100]. This condition is formally expressed as ∀i ∈ N, ∃s ∈ S such that h[i](s)̸ = f (s).
I think yes, this is an example of such an "i" and I would go so far as reclaiming that this is a pretty broad definition. Just saying that it is considered hallucinating if it makes something up that it was trained on (as opposed to something it wasn't trained on). I'm pretty confident the LLMs ingested a lot of English grammar books so I think it is fair to say that this was in the training.[0] https://cloud.google.com/discover/what-are-ai-hallucinations
My definition of "hallucination" is evidently not nearly as widespread as I had assumed.
I ran a Twitter poll about this earlier - https://twitter.com/simonw/status/1953565571934826787
All mistakes by models — ~145 votes
Fabricated facts — ~1,650 votes
Nonsensical output — ~145 votes
So 85% of people agreed with my preferred "fabricated facts" one (that's the best I could fit into the Twitter poll option character limit) but that means 15% had another definition in mind.
And sure, you could argue that "this sentence is in first person" also qualifies as a "fabricated fact" here.
If they were different things (objectively, not "in my opinion these things are different) then they'd be handled differently. Internally they are the exact same thing: wrong statistics, and are "solved" the same way. More training and more data.
Edit: even the "fabricated fact" definition is subjective. To me, the model saying "this is in first person" is it confidently presenting a wrong thing as fact.
It doesn't matter what you call it, the output was wrong. And it's not like something new and different is going on here vs whatever your definition of a hallucination is: in both cases the model predicted the wrong sequence of tokens in response to the prompt.
0 - Not faceplanting when trying to run
I usually use an agentic workflow and "hallucination" isn't the first word that comes to my mind when a model unloads a pile of error-ridden code slop for me to review. Despite it being entirely possible that hallucinating a non-existent parameter was what originally made it go off the rails and begin the classic loop of breaking things more with each attempt to fix it.
Whereas for AI autocomplete/suggestions, an invented method name or argument or whatever else clearly jumps out as a "hallucination" if you are familiar with what you're working on.