In general facts are not the answer to Hallucinations. You can't possibly have every fact for every situation. The true solution to Hallucinations is figuring out how to make a model say 'I don't know"
In general facts are not the answer to Hallucinations. You can't possibly have every fact for every situation. The true solution to Hallucinations is figuring out how to make a model say 'I don't know"
(Issue is, now some are convinced that people in general would do the same and just blurt out the feedforward output of their "internal neural network", as opposed to having built knowledge in a loop of critical evaluation.)
And why would that be? "Hallucination" means "erratic wandering", implying one is lost - similarly to "delirium" (maetaphor using the plough) and "error". Part of the idea is that of "instead of witnessing the correct, reporting the false" - a very ancient, traditional idea, and akin to the concept of "intelligence" (intus-legere).
"Confabulation" means locutor and interlocutor are talking, exchanging narrations.
but I am not sure - provisionally - that it can be a good idea to relate strictly human neurology to ANNs, if based on phenomena as opposed to structural issues. You do not have that problem when staying with natural language.
Still have to figure a measurement unit.
Second, it isn’t even necessarily better to have fewer lies if those few lies are more subtle. Plenty of propaganda works by twisting facts and using misleading statements. Perhaps the worst offenders won’t even have any outright falsehoods at all.
This has not been my experience. Did you create any benchmarks as a part of this project?
My startup has a product for lawyers that uses RAG to answer legal queries (https://lawlight.ai/). We have a disclaimer that "... (we) do not guarantee the accuracy of answers. You are responsible for reviewing the cited case law and drawing your own independent conclusions."
(This works within the specific context—lawyers are domain experts; and they are supposed to read through all cases they cite in court anyway.)
* I dislike the term "hallucinations." By definition LLMs hallucinate. It's just that much (or most) of the time, the hallucinations reflect reality.
There's a categorical difference between knowing a fact, and looking up a fact. When you know a fact you can recognize it in a situation where you wouldn't know to look it up, and you'd know to utilize it in a larger solution rather than simply parrot it when specifically asked about it.
Databases of facts have and will still have their place, but that is absolutely not the solution to LLM telling apart fact from truth. They have to innately have this in their model. I don't believe the nature of LLM is to hallucinate. It's instead a side effect of how we train them. We train them to guess, to be close, but not to be correct necessarily. And why is it a surprise that's precisely what they do?
Also LLM are too small in order to be accurate. They're tiny. GPT4 is roughly 40 times smaller than a human brain. And GPT4 is very large compared to GPT-3, and GPT-3 is very large compared to LLaMA 2.
We'll need for hardware to catch up so we can scale things up pragmatically and see what happens to their ability to grasp facts. But also architectural changes, of course.
Thoughout this comment you speak about LLMs as-if they're animals, or real physical objects. An LLM is a formal model which is just to generate a sequence of tokens maximally probabilistically consistent with a corpus of historical text.
A digital machine running a LLM program is a physical object which necessarily generates text based on "guessing" because that's the algorithm it's running. LLMs are "guessing algorithms", all of Machine Learning is -- it is dumb brute-force analysis of conditional probability.
> GPT4 is roughly 40 times smaller than a human brain
This doesn't make any sense. GPT4 is an abstract algorithm with no "size". The brain has 10^{big number} cells, and GPT4 can be specified with a single real number. Is that the comparison to make? No, both comparisons are incoherent.
A physical device running GPT4 can be given a "size", but it would again have nothing to do with a brain.
LLMs arent living things where we can "measure their size" and "train them to know, rather than to guess". They are just the equation, `max P(answer|propmt, historical_corpus)`
A machine running GPT4 is just an electrical device generating text according to the rule given above. There is no sense of "training it to do something other than guesswork", and no sense of "size"
What they are literally doing is guessing the next word, a word a time but doing it really really well and making statistically average output over a very large number of inputs.
There is no distinction between understanding "the" vs "a" and telling me 1+1=3. It is all token generation.
For what that concerns us here: LLMs will never learn to fact-check anything. They'll blindly regurgitate the facts they have been "taught", but never consider or evaluate "the paper cited for this fact on wikipedia is a bunch of bullshit".
Any attempt to use them to produce "facts" is ultimately just folly, in the same way Google's attempt to do so with it's search engine index is.
Nor do people, though! This is setting the bar way too high.
The whole point to having edited reference sources like "encyclopedias" is that so that we can rely on the expertise of the editors in lieu of having to develop the expertise ourselves[1].
No, an LLM that simply knows a priori (via prompt hacking) which sources are trustworthy would be absolutely comparable to the way an educated-but-non-expert human approaches sources.
[1] Which is a chicken and egg problem anyway. Everyone starts with edited reference sources as tutorial material. Quite frankly everyone starts learning with wikipedia.
No. If these things are claimed to be sources of truth, then the bar needs to be that high.
It is precisely because people don't fact-check that the bar has to be so high.
That's a strawman, though. No service, nor human, "claims to be a source of truth" in the kind of profound sense you seem to be using. It stops, everywhere, at "Wikipedia (or whatever) said it and I trust it".
The only way to get access to deeper expertise is to (1) BE an expert and (2) engage in an discussion with another.
GPT doesn't do math correctly but it also doesn't just memorize it.
What is most probable is not always what is most correct or most accurate.
Why is that necessary? Why have an LLM guess where the facts are?
Put all of that data in a place where it's normalized and ready to vector search.