Anthropomorphizing statistical learning is how you build a hype machine to cash out people with zero handle on the subject. See the comment below about "AI judges" and "true justice". Just like early electricity, all people see is magic.
Generative AI output is becoming inextricably associated with this word, and that's not a bad thing to keep people aware of.
_compared to what_, exactly. Compared to a google search? Compared to asking a random person? Compared to wikpedia? New York Times journalists?
Any of those things are wrong _very_ frequently. It's such an uninteresting thing to call out every time an AI is wrong, when it is right about things so frequently that people don't bother to notice how amazing it is that it gets anything correct about the world at all.
I've certainly made that class of error myself, when I assumed that something followed a similar pattern (like in math, or writing & grammar, or coding) when it actually didn't.
I've also doubled-down on those errors when I tried to double-check my work, believing myself to have misapplied some intermediate step rather than having taken an entirely wrong approach to begin with.
I think the "why" here is "why are we assuming this failure mode is unique to LLMs and deserves novel terminology".
We could find out 800 years from now that human brains really do work exactly the same as LLMs do, but it wouldn't change the fact that today LLMs and humans in practice regularly manifest their respective mistakes in very different ways.
For example I don't have to worry about you making sense but then turning on a dime saying things like "Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay. Okay." or "Subtitles by the Amara.org community", which are both examples of OpenAI hallucinations I encountered today.
We can call that type of stupidity its own word, different from the types of mistakes you described making, just like we have many different words to convey the concept of "wet".
The fact that LLMs are probably not just baby people give an even greater justification to use different terminology for them.
Synonyms thrive in our languge with sometimes a hair's breadth of difference in nuance so it's silly to let optimism about tech deny this.
I'm not sure how I feel about the term "hallucination" as it's applied to AI. Since you seem strongly opposed to it, let me ask you this long-winded half-question:
People understand computer things by creating analogies to the physical world - just look at the "Desktop" motif. "Folders" and "Files" too, for that matter. It seems to me that anthropomorphization would fit under that umbrella, though you may disagree. How do you feel about computer anthropomorphization in general? Is there something about "hallucination" that's particularly offensive?
Bugs, Defects and "not fit for production".
How about we stop with all the Nonsense around calling it "temperature" like it's a sick baby and call it RAND cause that's what it is.
The PT Barnum levels of bullshit around ML (see we have a term that isnt using artificial or intelligence) has gotten old. Sam Altman is the next Elizabeth Holmes.
</rant>
If I ask a software to write about a well known fact or historical event and it just makes stuff up, it's not simply hallucinating. It's defective.
The defect isn't in the software, but in people expecting these things to operate the way AIs in sci-fi do, or who believe that because they can produce coherent results in natural language, they must be sentient and self-aware.
I'm sure AI companies will get very good at explaining away these defects with various forms of "aCkShUaLlY" but when your marketing materials say you made a box that takes a prompt and answers it, and it answers incorrectly, what else is it than a defect?
Floating point math is inherently inaccurate, and no programmer using it would expect perfect precision and call it a defect not to get it. You have to understand how floating point works and take that inaccuracy into account. As a result there are some applications for which using floats is simply a bad idea. No one sane is doing real money calculations with floats.
The same goes for LLMs. Hallucination is fundamental to the model. We're going to have to realize that there are many tasks for which AI simply isn't well suited. And we're going to have to get over this persistent delusion that humans are categorically worse than AI at everything. A paralegal doing research would probably not simply fabricate cases and cites whole cloth. That's not how most humans work. Humans are capable of knowing when they don't know something, AI is not.
But we've decided, for whatever reason, that AI is perfectly trustworthy. That's going to keep biting us in the ass until we learn.
EDIT: Another way to put it: If I sold a calculator that claimed to do math, but in the fine print I said "Actually, it just makes up answers by some means we don't fully understand, and somehow most of the time it comes up with the right answer." That doesn't mean that incorrect answers are suddenly not defects.
So there's a lot of rigor on containing exactly how big your errors can be in floating point calculations, and that old intel bug made it (rarely) invalid.
Sugar coating the fact that it is defective (defined: imperfect or faulty.) isnt changing things.
Your explanation is correct, it's defective by design.
The fact that ‘correct’ outputs are treated as if they’re the product of an in-any-way-different process to the ‘hallucinated’ ones is the problem.
Also this particular context just makes it easier to notice, compared a 5000 word generated coherent-word-salad that equally wrong, but across the 5000 words.
Saying it hallucinated is just a tautology.
"Wargames" (1983): https://www.youtube.com/watch?v=71k7-dGhNFQ&t=4m8s