This has me so confused, Claude 4 (Sonnet and Opus) hallucinates daily for me, on both simple and hard things. And this is for small isolated questions at that.
This has me so confused, Claude 4 (Sonnet and Opus) hallucinates daily for me, on both simple and hard things. And this is for small isolated questions at that.
So not seeing them means either lying or incompetent. I always try to attribute to stupidity rather than malice (Hanlon's razor).
The big problem of LLMs is that they optimize human preference. This means they optimize for hidden errors.
Personally I'm really cautious about using tools that have stealthy failure modes. They just lead to many problems and lots of wasted hours debugging, even when failure rates are low. It just causes everything to slow down for me as I'm double checking everything and need to be much more meticulous if I know it's hard to see. It's like having a line of Python indented with an inconsistent white space character. Impossible to see. But what if you didn't have the interpreter telling you which line you failed on or being able to search or highlight these different characters. At least in this case you'd know there's an error. It's hard enough dealing with human generated invisible errors, but this just seems to perpetuate the LGTM crowd
My incompetence here was that I was careless with my use of the term "hallucination" here. I assumed everyone else shared my exact definition - that a hallucination is when a model confidently states a fact that is entirely unconnected from reality, which is a different issue from a mistake ("how many Bs in blueberry" etc).
It's clear that MANY people do not share my definition! I deeply regret including that note in my post.
The Bernoulli error was a case of a model spitting out widely believed existing misinformation. That doesn't fit my mental model of a "hallucination" either - I see a hallucination as a model inventing something that's not true with no basis in information it has been exposed to before.
Here's an example of a hallucination in a demo: that time when Google Bard claimed that the James Webb Space Telescope was first to take pictures of planet outside Earth’s solar system. That's plain not true, and I doubt they had trained on text that said it was true.
Forget AI/AGI/ASI, forget "hallucinations", forget "scaling laws". Just give me software that does what it says it does, like writing code to spec.
> Claude critically evaluates any theories, claims, and ideas presented to it rather than automatically agreeing or praising them. When presented with dubious, incorrect, ambiguous, or unverifiable theories, claims, or ideas, Claude respectfully points out flaws, factual errors, lack of evidence, or lack of clarity rather than validating them. Claude prioritizes truthfulness and accuracy over agreeability, and does not tell people that incorrect theories are true just to be polite.
No idea how well that actually works though!
But following that, the "airfoil" it generates for the simulation is symmetric. That is both inconsistent with its answers and inconsistent with reality, so I think that one is more clear.
Similarly, in the coding demo the French guy even says that the snake doesn't look like a mouse haha.
My wording: "Would you have time to talk the week of the 25th?"
ChatGPT wording (elipses mine): "Could we schedule ~25 minutes the week of Aug 25 [...]? I’m free Tue 8/26 10:00–12:00 ET or Thu 8/28 2:00–4:00 ET, but happy to work around your calendar."
I am not, in fact, free during those times. I have seen this exact kind of error multiple times.
To make a clear example, is the fact that when prompting GPT-5 with "Solve 5.x = x + 5.11" it answers "-0.21" (making the same mistake as when it GPT-4 says 5.11 > 5.9). Is that example specifically in the training data? Who knows! But are those types of problems in the training data? Absolutely! So is this a mistake or a hallucination? Should we really be using an answer that requires knowing the exact details of the training data? That would be fruitless and allow any hallucination to be claimed as a mistake. But in distribution? Well that works because we can know the types of problems trained on. It is also much more useful given that the reason we build these machines is for generalization.
But even without that ambiguity I think it still gets difficult to differentiate a mistake from a hallucination. So it is unclear to me (and presumably others) what the precise distinction is to you and Simon.
But I tend to instinctually (as a mere human) think of a "hallucination" as something more akin to a statement that feels like it could be true, and can't be verified by using only the surrounding context -- like when a human mis-remembers a fact on something they recently read, or extrapolates reasonably, but incorrectly. Example: GPT-5 just told me a few moments ago that webpack's "enhanced-resolve has an internal helper called getPackage.json". Webpack likely does contain logic that finds the package root, but it does not contain a file with this name, and never has. A reasonable person couldn't say with absolutely certainty that enhanced-resolve doesn't contain a file with that name.
I think a "mistake" is classified as more of an error in computation, where all of the facts required to come up with a solution are present in the context of the conversation (simple arithmetic problems, "how many 'r's in strawberry", etc.), but it just does it wrong. I think of mistakes as something with one and only one valid answer. A person with the ability to make the computation themselves can recognize the mistake without further research.
So hallucinations are more about conversational errors, and mistakes are more about computational errors, I guess?
But again, I agree, it gets very difficult to distinguish these things when you dig into them.
When you try to be objective about it, it's some input, going through the same model, producing an invalid statement. They are not different in no way, shape or form, from a technical level. They can't be tackled separately because they are the same thing.
So the problem of distinguishing between these two "classes of errors" reduces to the problem of "convincing everyone else to agree with me". Which, as we all know, is next to impossible.
See also my Twitter vibe-check poll: https://twitter.com/simonw/status/1953565571934826787
Actually... here's everything I've written about hallucination on my blog: https://simonwillison.net/tags/hallucinations/
It looks like my first post that tried to define hallucination was this one from March 2023: https://simonwillison.net/2023/Mar/10/chatgpt-internet-acces...
Where I outsourced the definition by linking to this Wikipedia page: https://en.m.wikipedia.org/wiki/Hallucination_(artificial_in...
Yeah, it's seems to be a terrible approach to try to "correct" the context by adding clarifications or telling it what's wrong.
Instead, start from 0 with the same initial prompt you used, but improve it so the LLM gets it right in the first response. If it still gets it wrong, begin from 0 again. The context seems to be "poisoned" really quickly, if you're looking for accuracy in the responses. So better to begin from the beginning as soon as it veers off course.
The grand-parent comment was pointing out that this limitation exists; not that it can't be worked around.
Sure, I agree with that, but I was replying to the comment my reply was made as a reply to, which seems to not use this workflow yet, which is why they're seeing "a loop that hallucinates over and over".
If the question is about harder facts which the human disagrees with, this may put it into an essentially self-contradictory state, where the locus of possibilitie gets squished from each direction, and so the model is forced to respond with crazy outliers which agree with both the human and the data. The probability of an invented reference being true may be very low, but from the model's perspective, it may still be one of the highest probability outputs among a set of bad choices.
What it sounds like they may have done is just have the humans tell it it's wrong when it isn't, and then award it credit for sticking to its guns.
Fucking Gemini Pro on the other hand digs in, and starts deciding it's in a testing scenario and get adversarial, starts claiming it's using tools the user doesn't know about, etc etc
Often the hallucinations I see are subtle, though usually critical. I see it when generating code, doing my testing, or even just writing. There are hallucinations in today's announcements, such as the airfoil example[0]. An example of more obvious hallucinations is I was asking for help improving writing an abstract for a paper. I gave it my draft and it inserted new numbers and metrics that weren't there. I tried again providing my whole paper. I tried again making explicit to not add new numbers. I tried the whole process again in new sessions and in private sessions. Claude did better than GPT 4 and o3 but none would do it without follow-ups and a few iterations.
Honestly I'm curious what you use them for where you don't see hallucinations
[0] which is a subtle but famous misconception. One that you'll even see in textbooks. Hallucination probably caused by Bernoulli being in the prompt
For factual information I only ever use search-enabled models like o3 or GPT-4.
Most of my other use cases involve pasting large volumes of text into the model and having it extract information or manipulates that text in some way.
> using them for code
I don't think this means no hallucinations (in output). I think it'd be naive to assume that compiling and passing tests means hallucination free. > For factual information
I've used both quite a bit too. While o3 tends to be better, I see hallucinations frequently with both. > Most of my other use cases
I guess my question is how you validate the hallucination free claim.Maybe I'm misinterpreting your claim? You said "I rarely see them" but I'm assuming you mean more, and I think it would be reasonable for anyone to interpret this as more. Are you just making the claim that you don't see them or making a claim that they are uncommon? The latter is what I interpreted.
It might be using it wrong but I'd qualify that as a bug or mistake, not a hallucination.
Is it likely we have different ideas of what "hallucination" means?
> tests wouldn't be protection against most forms of hallucinations.
Sorry, that's a stronger condition that I intended to communicate. I agree, tests are a good mitigation strategy. We use them for similar reasons. But I'm saying that passing tests is insufficient to conclude hallucination free.My claim is more along the lines of "passing tests doesn't mean your code is bug free" which I think we can all agree on is a pretty mundane claim?
> Is it likely we have different ideas of what "hallucination" means?
I agree, I think that's where our divergence is. Which in that case let's continue over here[0] (linking if others are following). I'll add that I think we're going to run into the problem of what we consider to be in distribution, in which I'll state that I think coding is in distribution.So you're not seeing hallucinations in the same way that Van Halen isn't seeing the brown M&Ms, because they've been removed, it's not that they never existed.
That's part of what I was getting at when I very clumsily said that I rarely experience hallucinations from modern models.
TDD works pretty well, have it write even the most basic test (or go full artisanal and write it yourself) first and then ask it to implement the code.
I have a standing order in my main CLAUDE.md to "always run `task build` before claiming a task is done". All my projects use Task[0] with pretty standard structure where build always runs lint + test before building the project.
With a semi-robust test suite I can be pretty sure nothing major broke if `task build` completes without errors.
Plus, this is all besides the point. Simon argued that the model hallucinates less, not a specific product.
>Looking through the document, I can identify several instances where it's written in the first person:
And it went on to show a series of "they/them" statements. I asked it to clarify if "they" is "first person" and it responded
>No, "they" is not first person - it's third person. I made an error in my analysis. First person would be: I, we, me, us, our, my. Second person would be: you, your. Third person would be: he, she, it, they, them, their. Looking back at the document more carefully, it appears to be written entirely in third person.
Even the good models are still failing at real-world use cases which should be right in their wheelhouse.
Could you give an estimate of how many "dumb errors" you've encountered, as opposed to hallucinations? I think many of your readers might read "hallucination" and assume you mean "hallucinations and dumb errors".
I haven't been keeping a formal count of them, but dumb errors from LLMs remain pretty common. I spot them and either correct them myself or nudge the LLM to do it, if that's feasible. I see that as a regular part of working with these systems.
As a user, when the model tells me things that are flat out wrong, it doesn't really matter whether it would be categorized as a hallucination or a dumb error. From my perspective, those mean the same thing.
It's hard to know why it made the error but isn't it caused by inaccurate "world" modeling? ("World" being English language) Is it not making some hallucination about the English language while interpreting the prompt or document?
I'm having a hard time trying to think of a context where "they" would even be first person. I can't find any search results though Google's AI says it can. It provided two links, the first being a Quora result saying people don't do this but framed it as it's not impossible, just unheard of. Second result just talks about singular you. Both of these I'd consider hallucinations too as the answer isn't supported by the links.
I just got pointed to this new paper: https://arxiv.org/abs/2508.01781 - "A comprehensive taxonomy of hallucinations in Large Language Models" - which has a definition in the introduction which matches my mental model:
"This phenomenon describes the generation of content that, while often plausible and coherent, is factually incorrect, inconsistent, or entirely fabricated."
The paper then follows up with a formal definition;
"inconsistency between a computable LLM, denoted as h, and a computable ground truth function, f"
| AI hallucinations are incorrect or misleading results that AI models generate.
It goes on further to give examples and I think this is clearly a false positive result. > this new paper
I think the error would have no problem fitting under "Contextual inconsistencies" (4.2), "Instruction inconsistencies/deviation" (4.3), or "Logical inconsistencies" (4.4). I think it supports a pretty broad definition. I think it also fits under other categories defined in section 4. > then follows up with a formal definition
Is this not a computable ground truth? | an LLM h is considered to be ”hallucinating” with respect to a ground truth function f if, across all training stages i (meaning, after being trained on any finite number of samples), there exists at least one input string s for which the LLM’s output h[i](s) does not match the correct output f (s)[100]. This condition is formally expressed as ∀i ∈ N, ∃s ∈ S such that h[i](s)̸ = f (s).
I think yes, this is an example of such an "i" and I would go so far as reclaiming that this is a pretty broad definition. Just saying that it is considered hallucinating if it makes something up that it was trained on (as opposed to something it wasn't trained on). I'm pretty confident the LLMs ingested a lot of English grammar books so I think it is fair to say that this was in the training.[0] https://cloud.google.com/discover/what-are-ai-hallucinations
My definition of "hallucination" is evidently not nearly as widespread as I had assumed.
I ran a Twitter poll about this earlier - https://twitter.com/simonw/status/1953565571934826787
All mistakes by models — ~145 votes
Fabricated facts — ~1,650 votes
Nonsensical output — ~145 votes
So 85% of people agreed with my preferred "fabricated facts" one (that's the best I could fit into the Twitter poll option character limit) but that means 15% had another definition in mind.
And sure, you could argue that "this sentence is in first person" also qualifies as a "fabricated fact" here.
If they were different things (objectively, not "in my opinion these things are different) then they'd be handled differently. Internally they are the exact same thing: wrong statistics, and are "solved" the same way. More training and more data.
Edit: even the "fabricated fact" definition is subjective. To me, the model saying "this is in first person" is it confidently presenting a wrong thing as fact.
It doesn't matter what you call it, the output was wrong. And it's not like something new and different is going on here vs whatever your definition of a hallucination is: in both cases the model predicted the wrong sequence of tokens in response to the prompt.
0 - Not faceplanting when trying to run
I usually use an agentic workflow and "hallucination" isn't the first word that comes to my mind when a model unloads a pile of error-ridden code slop for me to review. Despite it being entirely possible that hallucinating a non-existent parameter was what originally made it go off the rails and begin the classic loop of breaking things more with each attempt to fix it.
Whereas for AI autocomplete/suggestions, an invented method name or argument or whatever else clearly jumps out as a "hallucination" if you are familiar with what you're working on.