The behavior of text models is similar enough that the wording stuck, and it's not all that bad.
Rather than just an imperfect technology as we have here.
Many people object to the term enshittification for foul-mouthing reasons but I think it covers it very well because the principle it covers is itself so very nasty. But that's not at all the case here.
But there's nothing at all different about what the model is doing between these cases -- the models are hallucinating all the time and have no ability to assess when they are hallucinating "right" or "wrong" or useful/non-useful output in any meaningful way.
: to fill in gaps in memory by fabrication
> In psychology, confabulation is a memory error consisting of the production of fabricated, distorted, or misinterpreted memories about oneself or the world.
It’s more about coming up with a plausible explanation in the absence of a readily-available one.
They simply do not behave like humans of sound minds, and "hallucinations" conveys that in a way that "confabulations" or even "bullshit" does not. (Though "bullshit" isn't bad either.)
I think what calling the times they get things wrong hallucinations is largely an advertising trick. So that they can sort of fit the LLMs into how all IT is sometimes “wonky” and sell their fundamentally flawed technology more easily. I also think it works extremely well.
It is the memory pathways leading them astray. It could be thought of a memory system that at certain point any longer can't be fully sure if whatever connections they have are from actually being trained or it or created accidentally.
I suppose so, in the sense that someone could simply be lying about pink elephants instead of seeing them. However it's hard to argue that the machine knows the "right" answer and is (intelligently?) deceiving us.
> It is the memory pathways leading them astray.
I don't think it's a "memory" issue as much as a "they don't operate the way we like to think they do" issue.
Suppose a human is asked to describe different paintings on the wall of an art gallery. Sometimes their statements appear valid and you nod along, and sometimes the statements are so wrong that it alarms you, because "this person is hallucinating."
Now consider how the entire situation is flipped by finding out one additional fact... They're actually totally blind.
Is it a lie? Is it a hallucination? Does it matter? Either way you must dramatically re-evaluate what their "good" outputs really mean and whether they can be used.
Like you ask me for a birthdate of some obscure political figure from history? I'm going to try to feel out what period in history the name might feel like to me and just make my best guess based on that, then say some random year and a birthdate. It just has the lowest odds of being beaten. Was I hallucinating? No, I was just trying to not get beaten.
This is transparently wrong. It gets so many things right in a response that the few things it gets wrong are tremendously frustrating. I think people underestimate how much correct "knowledge about the world" is expressed in a typical chat gpt response and focus only on the parts that are incorrect.
If it were wrong about _everything_ at rates no better than chance, we wouldn't even be having this conversation because nobody would be using them.
Edit: It retrospect, perhaps a better analogy would involve gasoline, as its explosive nature is what's being actively being exploited in normal use.
"Lastly, I want to reassure investors and members of the press that we take these concerns very seriously: The Ford Pinto-II will only contain only normal and stable gasoline, and not the rare and unusual burning kind, which is merely a temporary hurdle in this highly dynamic and explos--er--fast growing field."
And I couldn't find a single one of my friends who hadn't experienced "phantom vibration syndrome".
Both I'd say are "Hallucinations", without any real negative connotation.
LLMs don't do it because they are out of their right mind. They do it because every single answer they say is invented caring only about form, and not correctness.
But yeah, that ship has already sailed.
But "hallucination" was already (before LLMs) being used in a figurative sense, i.e. for abstract ideas that are made up out of nothing. The same is also true of other words that were originally visual, like "illusion" and "mirage".
Humans hallucinate. Programs have bugs.
It's inherent to how LLMs work and is expected although undesired behaviour.