What does it mean when an LLM “hallucinates” & why do LLMs hallucinate?
labelbox.com
labelbox.com
Stopped reading here. A model has no objective. It has only output.
The only objective is on the part of "AI" marketeers and it is there that the euphemism "hallucination" originates.
Is creating output not an objective?
But an objective function is a core element of ML models, no? And LLMs have objective functions, so far as I'm aware.
In star Trek there's the prime directive, the enterprise crew follows it because it's a core requirement of the federation.
Just because the LLM has an objective to answer queries for example does not mean it chose it. or has the intent itself as sapient being would.
They (chatgpt et al) do move closer to “clear and coherent” by adding extra layers such as a beam search on LLM outputs. Good to remember that ChatGPT et al are products, not bare metal LLMs.
The "AI" marketeers' use of "hallucinate", "objective" and "predict" does precisely that.
This article isn't doing that, though. What it is calling the model objective is "producing clear and coherent text".
> Whether "producing clear and coherent text" is a good characterisation of a "predict the next token" loss function is another question though.
I agree. But lets us be clear that prediction too is a misnomer. Simply generating a word coherent with prior words is not prediction.
I'm not buying :)
So LLMs hallucinate because that's what their architecture demands of them for the purpose of generating outputs. It's really that simple.
(But seriously, "confabulate" is much better. The thing I don't like about "hallucinate" is that it gets the agency wrong. Hallucinations are something that happens TO people, that they recover from.)
chatGPT is like a subordinate afraid of looking bad or being chastised for not knowing something it tells you what it thinks you want to hear. bullshitting is what I call this too.
If you use retrospection, and tell it to only give answers that are vetted or true or that it should retrospect on the validity of an answer because they could be true but only in so much as the LLMs cut-off date, for instance using outdated docs.
i personally think everyone working on ai should band together and create an open source framework or tool that's essentially an LLM verification API to verify truthyness.