LLMs confabulate not hallucinate
beren.io
beren.io
She can’t read, write, handle money, but confabulates engaging and convincing stories about receiving mail, haggling with her bank, her boyfriend. She speaks with proper grammar, intonation, etc, but doesn’t demonstrate any understanding of what she talks about beyond the words. She doesn’t have a bank account or a boyfriend.
Trying to find more and I can’t find anything about her on the internet, I would have to look at the references. Maybe it’s questionable.
Anyway, I like to think about what the hell we even mean when we say we “know something”, and where we should actually put the bar for AI. Where is the line between confabulation and knowledge? Is knowledge just confabulation that happens to be correct?
I feel like it’s possible that Denys could feel that she “knows” what she is saying, but is simply trapped in, with a working knowledge of the world and is just unable to translate it into actions, and words are just her only degree of freedom.
I don’t remember where I was going with this, but I agree that “confabulation” is a much better term to use, but we also shouldn’t discount what it means to be able to confabulate.
The tactic is simply to talk relevantly until the listener fools themselves into thinking they understand you, except that "understanding" will be an illusion and 100% manager's own analytical effort. And it's honestly embarrassing how often and well that works. Quite a few of those guys were promoted first, too.
I'm pretty sure our upcoming future with AI will be like that quote from forest ranger at Yosemite National Park, on why it is hard to design the perfect garbage bin to keep bears from breaking into it: "There is a considerable overlap between the intelligence of the smartest bears and the dumbest tourists."
The whole idea of a criteria to separate humans from AI is a misconception. I think any human who confuses smooth talking with making sense will be ruthlessly exploited in relatively near future. There are humans who already do this successfully to other humans, what do you think will happen when they get to automate their efforts.
Ads already ruthlessly exploit these people, taking their money in exchange for garbage products. LLM might make this worse, but we have already automated displaying targeted manipulative messages via ad networks. People who can handle manipulative messages from ads should be mostly resistant to manipulative messaging from bots.
Not necessarily. LLMs at the model of GPT-4, and eventually above, affect the economics of advertising. Currently, there's only so much sense in optimizing your ads before it's cheaper to just spam more of it (er, "increase exposure"). LLMs will make it cheaper to optimize the ads further (i.e. get more bang for the same buck), so I expect the baseline manipulative effect of most ads to increase.
It's one of the end-game consequences of the change in advertising economics I mentioned: with costs of future AI models going down, eventually companies may be able to afford to generate individually personalized and highly persuasive ads on the fly. A personal dedicated con artist for everyone, indeed.
In fact, the odds are extremely high that you, me, and everyone else here has been taken in by some kind of ad or ad-like rhetoric and internalized it as just being normal world-view.
If there exist those who haven't even been fooled in such a way, they must certainly have the self-awareness to nevertheless entertain doubts on the matter.
Versus the last thousand years?
Just wait until you learn about organized religion it’ll blow your mind.
That gets super complicated, because arguably religion is less about its content, and more about the community. The hunger for connection isn't going away.
In a sense, religion is the perfect intersubjective phenomenon - the actual beliefs don't objectively matter and can be arbitrary[0], what's important is that everyone in your community subscribes to the same arbitrary set of beliefs. Traditions and proclamations are effective proxies for community cohesion.
(Arguably, national propaganda is the same story, just played on fast-forward.)
Point being, religion is much more than being smooth talked into believing nonsense. It digs roots deep into how humans form societies.
--
[0] - Mostly. Some beliefs are more sticky or self-reinforcing in the real world than others. Religious beliefs undergo natural selection too.
All memes do - that's why their name is based on genes. What's interesting is discovering clusters of memes that are self-reinforcing and are passed down consistently, without deviating too much from the original (while still allowing changes), over a long period of time. (EDIT: it's like finding a spaceship or oscillator in Conway's Game of Life) Arguably, religions built around such clusters have an easier time getting adherents and can get resurrected tens or hundreds of years after being nominally defeated or purged. So I don't think the "actual beliefs (mostly) don't objectively matter" characterization is warranted. I think they do matter a lot from the evolutionary standpoint; not so much for any of the current adherents, but for their grandchildren and later descendants.
To what extent the memetic content carries the institution versus how much an established institution contributes to propagating the content is unclear to me. There's definitely a relationship there, a complex interplay between vested interests, provided services, and people's needs and expectations. As you said, religion is much more than just a set of "nonsense" people get talked into. This makes it hard to discern the real utility of any particular nonsense for ensuring the religion survival.
Possibly relevant:
LLMs confabulate not hallucinate
https://www.beren.io/2023-03-19-LLMs-confabulate-not-halluci...
But creating a new cult (let's call them all "cults") is rare. It takes a person with a unique combination of skills and talents, I bet most of us can't start a religion or succeed at selling snake oil at will. But with help of LLMs...
Buddy, they already exist, and they're carrying MBAs, Teaching Degrees, and working the Stock Market. Anecdotally, one of the few bona-fide psychopaths I met face to face submitted a thesis that everyone in the school system bent over backwards to say was amazing, on actual review had submitted a complete word-salad.
Instead I would say they DREAM. Dreams follow some logic but basically they are disjoint from reality. They involve same characters as our real-life experiences but what those characters do in a dream is not based on reality.
So one could think that ALL LLMs ever do is dream but much of the time their dreams are dreams which feel very real.
Our dreams are basically an LLM based on our real experiences but recombined in arbitrary ways, still retaining some logical structure but now detached from reality because when you recombine different experiences including our thoughts from real life, the result cannot really reflect reality very well. LLMs are better in this respect than our dreams, but sometimes their "dreams" get really obviously detached from reality.
Or if you want to use the term "hallucinate" with LLMs then fine but its more like they hallucinate all the time, it's just that sometimes, even often, their hallucinations seem to agree with the reality so well that we cannot tell the difference.
LLMs do not describe reality because they cannot experience reality, all they can give us is an "average description" created from many existing descriptions.
One way of looking at it is to say that all LLMs tell us is hearsay.
Hallucinations are much better described as noise. Its noise that lowers the quality of the signal. its just unlike white noise, this is in the form of coherent sentences/imagery.
Its not rationalising, its statistically choosing the next token. Also, generally the LLM doesn't get defensive (although its not always the case.)
Its not like a dementia patient not finding their keys and gradually convincing themselves that someone stole them instead. Because the alternative, that they are loosing cognitive faculties, is too horrific to acknowledge.
This is why its so annoyingly problematic. too much humanising, and not enough witnessing someone with dementia "confabulating"
It is noise, and nothing else.
We know how the LLMs were trained, so re-stating it doesn't help in anything. The point is that after the LLMs are trained they behave in certain ways and it can be helpful to say something about how it behaves.
For example, we can talk about how a linear regression can or cannot capture the causal effect of X1 over Y. "It's not capturing the causal effect. It's just minimizing the squared error" is an unhelpful statement.
Unprompted meaning we have shut down almost all sensory input (there seems to be some residual hearing and proprioception).
and, FWIW, confabulation is a lifetime terror for me -- with lucid dreams and https://www.psychologytoday.com/us/blog/dream-catcher/201909...
i fully wonder if LLMs are just trying to fill in gaps in direct knowledge, and seeking the closest parable they can find. So definitely confabulation rather than hallucination
And hearsay is usually a lot better than LLM output. I wouldn't use that word.
GPT-4 logits calibration pre RLHF - [https://imgur.com/a/3gYel9r](https://imgur.com/a/3gYel9r)
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - [https://arxiv.org/abs/2305.14975](https://arxiv.org/abs/2305...
Teaching Models to Express Their Uncertainty in Words - [https://arxiv.org/abs/2205.14334](https://arxiv.org/abs/2205...
Language Models (Mostly) Know What They Know - [https://arxiv.org/abs/2207.05221](https://arxiv.org/abs/2207...
It seems very much like hallucinations aren't a randomness or representation issue. The computation already knows. It just doesn't care about telling you this
The amount of BS we feed each other online, maybe it just thinks we'd prefer something made up than "I don't know"
For the record I wasn't saying the parent comment was BS :)
Matrix multiplications cannot feel anything so of course it doesn't care. Humans care. Language models do not.
So not unlike early LLM chat in areas where the model had not a lot of data about.
Here is another interview, this one between a fourteen-year-old girl called Denyse and the late psycholinguist Richard Cromer; the interview was transcribed and analyzed by Cromer's colleague Sigrid Lipka.
> I like opening cards. I had a pile of post this morning and not one of them was a Christmas card. A bank statement I got this morning!
< [A bank statement? I hope it was good news.]
> No it wasn't good news.
< [Sounds like mine.]
> I hate . . . , My mum works over at the, over on the ward and she said "not another bank statement." I said "it's the second one in two days." And she said "Do you want me to go to the bank for you at lunchtime?" and I went "No, I'll go this time and explain it myself." I tell you what, my bank are awful. They've lost my bank book, you see, and I can't find it anywhere. I belong to the TSB Bank and I'm thinking of changing my bank 'cause they're so awful. They keep, they keep losing . . . [someone comes in to bring some tea] Oh, isn't that nice.
< [Uhm. Very good.]
> They've got the habit of doing that. They lose, they've lost my bank book twice, in a month, and I think I'll scream. My mum went yesterday to the bank for me. She said "They've lost your bank book again." I went "Can I scream?" and I went, she went "Yes, go on." So I hollered. But it is annoying when they do things like that. TSB, Trustees aren't. . . uh the best ones to be with actually. They're hopeless.
I have seen Denyse on videotape, and she comes across as a loquacious, sophisticated conversationalist—all the more so, to American ears, because of her refined British accent. (My bank are awful, by the way, is grammatical in British, though not American, English.) It comes as a surprise to learn that the events she relates so earnestly are figments of her imagination. Denyse has no bank account, so she could not have received any statements in the mail, nor could her bank have lost her bankbook. Though she would talk about a joint bank account she shared with her boyfriend, she had no boyfriend, and obviously had only the most tenuous grasp of the concept "joint bank account" because she complained about the boyfriend taking money out of her side of the account. In other conversations Denyse would engage her listeners with lively tales about the wedding of her sister, her holiday in Scotland with a boy named Danny, and a happy airport reunion with a long-estranged father. But Denyse's sister is unmarried, Denyse has never been to Scotland, she does not know anyone named Danny, and her father has never been away for any length of time. In fact, Denyse is severely retarded. She never learned to read or write and cannot handle money or any of the other demands of everyday functioning.
Denyse was born with spina bifida ("split spine"), a malformation of the vertebrae that leaves the spinal cord unprotected. Spina bifida often results in hydrocephalus, an increase in pressure in the cerebrospinal fluid filling the ventricles (large cavities) of the brain, distending the brain from within. For reasons no one understands, hydrocephalic children occasionally end up like Denyse, significantly retarded but with unimpaired—indeed, overdeveloped—language skills. (Perhaps the ballooning ventricles crush much of the brain tissue necessary for everyday intelligence but leave intact some other portions that can develop language circuitry.) The various technical terms for the condition include "cocktail party conversation," "chatterbox syndrome," and "blathering."
Source: http://f.javier.io/rep/books/The-Language-Instinct-How-the-M...Cromer, R.F. 1991. "Language and thought in normal and handicapped children." Cambridge, MA: Blackwell
https://www.amazon.com/Language-Handicapped-Children-Cogniti...
> My bank are awful, by the way, is grammatical in British, though not American, English.
How would someone phrase this in idiomatic American English?
American English views organizations as a singular entity, not so much a group of individuals.
In all seriousness, this feels weirdly similar to listening to Donald Trump.
I think it's the "jumping from one subject to the next without a broader point" thing that made me think of Donald Trump.
The shunt isn’t something we have to think about very often, but it’s truly a remarkable piece of hardware that makes it so she can live a fairly normal life.
For anyone wondering:
https://en.wikipedia.org/wiki/Cerebral_shunt
> A cerebral shunt is a device permanently implanted inside the head and body to drain excess fluid away from the brain.
And that that content is exactly the sort of thing that would make a successful viral TikTok video.
Much like how LLMs will gladly and emphatically inform you that they're sentient beings with real feelings if given the right leading questions.
For instance, before transformers, when we finetuned ESRGAN for specific styles/functions, we used to say it would "hallucinate" in detail that couldn't possibly be in the original image pixels. GenAI wasn't much of a thing back then, and we never thought of the small GAN models as "remembering" things even though thats common language in transformer models. It felt more like we were teaching the GAN a job or style, not encoding memory.
And I think it fit! "Hallucinating" detail into a blurry pixel blob feels like a more accurate analogy to the human condition, and we weren't at the point where it was "confabulating" and generating big objects out of the blue like a diffusion model can.
See https://blog.research.google/2015/06/inceptionism-going-deep... and https://distill.pub/2017/feature-visualization/.
I would differentiate these dreamy generations (i.e., things that we don't see ordinarily) from generations that are incorrect. Sure, there is overlap.
https://neurosciencenews.com/peripheral-vision-brain-illusio...
see the abstract of this paper for example Face hallucination using Bayesian global estimation and local basis selection
> ...hallucinating the high-resolution detail of a low-resolution input face image.
while that's a different domain than text, in a quick search, the oldest reference i found to "hallucination" using vector-based word representations is from 1996 (there could easily be older references, i just didn't bother searching further):
Text Databases and Information Retrieval https://dl.acm.org/doi/pdf/10.1145/234313.234371
> The fixed-length vector is especially useful for parallel and hardware systems, but this method can sometimes hallucinate words that don't actually appear in the original document.
edit: including a link to the face hallucination paper i mentioned (there are lots more) https://www.semanticscholar.org/paper/Face-hallucination-usi...
Both words imply that there's something erroneous or anomalous happening when an LLM emits incorrect facts, as if the model can be in either a correct or incorrect state. But from a system perspective, there's no difference between the two scenarios: the model is consistently generating the most probable output based on its weights and context. It doesn't sometimes confabulate... it always confabulates.
The concept of "truth" or "falsehood" are something purely external to the model that humans bring when evaluating the output. They aren't part of the model itself.
I recall seeing a post on here that roughly just said, "stop saying hallucinate", which sure, is technically correct ("the best kind of correct"), but not super helpful. Contrast that with this post, which offers a constructive suggestion: replace it with "confabulate".
I agree that C is an improvement over H, for different reasons from the author. Here's why: From a public comms standpoint, the big problem with saying "hallucinate" is that it leads to misconceptions, because people think they know what "hallucinate" means, and may anthropomorphize their pop psych grasp of H onto the LLM.
Just like H, "C" seems to have a technical meaning in the psych literature, but it's far less in common parlance than "H", so lay readers have no fixed idea of what exactly "C" means. Nobody goes around experimenting with "confabulagens" to seek a spiritual high. Also, "confabulate" retains a non-clinical meaning that's similarly obscure, from long before the psychologists adopted it, and it's pretty evocative of what's really going on here (e.g. via related words like "fable" and "fib"). So "confabulate" seems ripe for adoption to describe this specific computational phenomenon, as we've done to so many other mundane words in the past (deprecate, and so forth).
If the psychiatric "C" phenomenon is also a better analogy for LLM behavior, as OP claims (and some posters dispute?), so much the better. But my main reason for liking C, is that it's less likely than H to make people unduly confident that they understand it.
(But even if the community came to a consensus at this late stage that "C" is a better term, the subsequent problem of updating every stale paper and post that uses the old term would run us up against another doozy: cache invalidation.)
Hallucinate generally implies experiencing sensory inputs which aren't there, but having a rational reaction to them - in the popular psyche though, people who hallucinate are still "crazy" - unfairly so, because you can function normally with hallucinations under quite a number of circumstances (i.e. I remember someone saying that they realized that if they saw people without faces, it was fine because they weren't really there - apparently quite common, since face recognition is a different part of the brain).
Basically: reacting to invalid sensory inputs implies rational behavior, but to non-real inputs.
Confabulation on the other hand is different because it's closer to "trying to rationalize an autonomous behavior without knowledge of the sensory inputs causing it" (maybe, it's a bizarre phenomenon). In split-brain patients in lab settings, it manifests as being unable to answer a question about why you're having a physiological response to imagery which is being shown only to the right-side of the brain, whereas language is processed on the left - but rather then be confused, people will apparently make something up that they're unaware is actually a "lie" (in quotes because, well, they're not lying - as far as we know there's no intent or even knowledge that is a lie).
Where this leaves us with LLMs I don't really know, because it feels like an imperfect descriptor: except perhaps with the context that LLMs are fairly limited networks, and so to some extent the whole "being wrong" phenonmenon is in fact just a failure of the attention mechanism to be able to draw the right data together.
No.
Firstly its just using a far more obscure term, moreover its common translation is "honest lies" which is an oxymoron.
Hallucinate is an accessible term that simply describes the outcome of what the user sees.
if you want a better term diving deep into obscure psych language is not the right way, it just confers an air of fanciness to what is either a feature, or a massive drawback of LLMs (depending on context)
Better words might be: flights of fancy, bullshit, imagined, illusion, rubbish, noise.
Indeed, the Latin verb "confabulor has the meaning of "discuss, converse" [0], not specifically "inventing things", and Romance languages already have words derived in sound and meaning from this classic Latin sense. For instance French already has "confabuler" meaning "speak familiarly with someone" [1]. Thus in Romance languages, this proposed usage of "confabulate" imported from the psychiatric word could be understood poorly due to preexisting similar-sounding words that do not include the notion of "making things up".
Just proposing thoughts here, not advocating for a final decision. Maybe it's still good to adopt this wording!
[0] : https://en.m.wiktionary.org/wiki/confabulor#Latin [1] : https://www.dictionnaire-academie.fr/article/A9C3479
That's why cardinals are trapped in a "confab"
edit: it is: https://www.etymonline.com/word/confabulate -- although both the verb fabulari and the noun fabula ultimately come from the proto root bha-, which is also the root for falar in Portuguese and hablar in Spanish (and presumably parlez in French?)
The Italian verb parlare has a similar etymology.
I think of LLMs as world-class bullshitters. They often hold forth with complete confidence on topics they have little or no understanding of, making up what they don't know, and appearing to be experts. Except to somebody who actually knows something.
Like that blowhard friend who has an opinion on anything.
Yes. I just keep getting reminded of the episode of TNG when Data learns to smalltalk by mimicking an infamous blowhard from Starfleet [1].
(Genuine) Question then: Does "No that's wrong"... every do something? Is there any reason to do so other than having fun arguing? I.e. if it gives wrong answer initially, and the user knows it's the wrong answer... what's the goal/desired outcome in telling it it's wrong and asking it again? Does that goal every get accomplished? What is the confidence in the second answer?
From a less black box perspective, is there any "ELI5" or at least "ELI45" explanation of why it might ever work? :)
The answer to "What are LLMs but humans with extreme amnesia and no central coherence?" is that they are not like humans at all, and only resemble them superficially due to our anthropomorphic tendency to infer a mind upon seeing fluent and roughly correspondent text.
Alternatively, we overestimate human abilities, and imbue our thoughts with deep meaning and significance, simply as a way to feel more important than we actually are. Ultimately, we may rely on statistical prediction quite a bit as well.
Obviously, LLMs aren't at the level of human cognition, but the impressive results they do produce, given their obvious simplicity, hints that intelligence may not be as hard to reproduce in silicon as we have previously assumed.
People said this about calculators as well, or initial chess AI as well. "If computers can do this, then the remaining parts of intelligence is probably not hard to reproduce!".
So we already know many humans massively underestimate the depth of human intelligence. It happened many times before and will continue to happen every time we make computers do anything new.
People also said that chess computers would never beat the best humans. And then when that happened they said, "well they're just using sheer brute force, they will never produce beautiful chess". And then Alpha Zero came along, and showed that prediction was wrong as well.
Personally, I believe that silicon intelligence will evolve a lot faster than ours did. But regardless, I've yet to hear any convincing argument why it can't happen eventually.
That is a strawman, basically nobody argues this.
> People also said that chess computers would never beat the best humans
How many did say this? I see way more overly optimistic "humans are stupid" messages than I see "computers will never beat humans" messages. The argument is mostly between "humans are obsolete in a decade at most" versus "we would be lucky to see general intelligence in our lifetime", not that it it is impossible.
I don't know about chess, though earlier computers didn't have the power, so people could've reasonably said that.
But we certainly saw it with Go. Many people said that a computer could beat a person at chess or checkers, because they were simple, limited games, but Go was too complicated, too many possible combinations. A computer could never beat a Go champion. And here we are, with multiple generations of AlphaGo pushing the boundaries. They might use a couple thousand CPUs and a few hundred GPUs, but they can do the impossible thing.
* Using the algorithms they used to beat humans at chess.
I were there during those discussions, the AI optimists argued that since computers could beat humans at chess GAI would soon be here. Then it is reasonable to argue that even a simple game like GO is impossible to solve the way we solved chess, so there needs to be more fundamental breakthroughs before GAI can happen.
That same thing is playing out today, and the AI proponents are still misunderstanding the other sides argument. The argument isn't "computers can never do this", the argument is "the algorithms and methods we use today can't do this, so there is no clear path to intelligence from where we are".
Some people do. Some people believe that intelligence is much more fundamental to the universe, and not capable of being reproduced mechanistically.
But if you grant that it is possible, all we're talking about is when; which is anyone's guess. It could be tomorrow, it could take a million years. I'm not pretending to know, just offering a different perspective than the original comment.
I mean, people do-- even smart people like Penrose.
Some people think we can't get to "real" intelligence through Turing-equivalents. Some think we can just tack a little more on an LLM and be there. It's probably something between, but the past-me who thought AGI was definitely >40 years away is much less convinced of this now.
Could it happen in the next 5 years? Proooooobably not, but is there a significant chance? Yes.
To it, you've basically said: "well in that case defend the times some other people have been incorrect if you're sure this time".
Why should they? The OP might be wrong, but whether they are or aren't has nothing to do with whether some unrelated, half-remembered third parties were right or wrong when mispredicting something else. You haven't addressed the content of the OP's actual argument at all, nor established a basis on which it may share similarity with other mispredictions.
Their argument was "Alternatively, we overestimate human abilities". Since he made that argument it is relevant to see if humans overestimated or underestimated human abilities in the past. My argument is that there is a large tendency for smart people and researchers to underestimate human intelligence, to assume it is easy to solve, we have seen that happen many times in the past.
So since smart humans often underestimate human abilities, the "we overestimate human abilities" theory is probably not correct.
We have an enormous history of vastly over-estimating the intelligence needed for some behaviors to manifest, with the more accurate conclusion being that actually we simply don't understand the nature of intelligence very well.
So yeah, underestimating animal intelligence is a large part of it, I agree with you.
If I ask you to envision a green triangle and a red square next to each other, and then swap the shapes but keep the colors in the same locations, and answer what color is the triangle now, you say the triangle is red, but you do so because you envisioned the triangle swapping places and did the mental steps etc.
An LLM if even answering correctly, is statistically answering based on billions of lines of text + rlhf and all of this, I highly doubt there is a mental model of the world, but rather a large set of constraints in the probabilities which leads to the resulting answer. The reasoning ability is a secondary effect of the probabilities which is why it's hard to make it so every probability is correct for every answer I think.
And regarding OP about hallucinating vs confabulating. To me hallucinating is a fine word for it, because it is filling in a gap or there aren't enough constraints in the model/data/tuning to account for that specific answer that it gave that was incorrect. Hence it "hallucinates" something in the gap. The real power of LLM's is that it seems to accumulate these 'constraints' (generalization), so that with the right model, it should be able to answer more and more prompts correctly, which is kind of amazing.
Confabulation works too but is a little more high level IMO.
The triangle example doesn't prove it because our predictive model could also say "hey, things don't just change color and shape like that, they need to move". It's similar to how an LLM can be more accurate with math when asked to step through and reason through it's logic - by stepping through the individual steps it can create a larger system.
The thing that made me question if a "mental model" is at the base of human cognition - people who do those memory competitions, the clear winning strategy is the memory palace, or imagining walking through a house where each room is another number - they have to build step by step memory, it's not like an SQL database where they can just SELECT from random. Another one was the insight from GTD that if you remember 7 things to do, you're always repeating those 7 things to yourself to keep them in active memory.
There's a strong argument that the mental model is derivative of a predictive model in the human brain, and we can just appear to have a mental model since we have an internal dialogue that runs so fast in the background that even we rarely recognize it. (anyone who has kept a steady meditation or similar practice should be familiar with it)
That's the difference between humans and llm's imo. We can generalize any "computation" we have to any other "desired output" we want, by thinking about it, while llm's aren't at least not now, so general that they use low level representations of all the 'objects' we can prompt about. Like humans can reason about the objects and things in our mind almost infinitely and recursively while also retaining all the physical realities and facts of those objects, while an llm is limited in this regard. Doesn't mean it can't in theory, there is some weird generalization going on as far as I can tell, but it feels like it's going to need a lot more data or something to do it.
It is akin to a backronym "well that sounds similar so is the same" while ignoring the real differences.
But humans clearly do that. It's not all we do, but we clearly do that. And given the impressive results that LLMs produce, my guess is it or something like it, will play a significant role in the AGI of the future.
What I do now, and how I think about things is being affected in ways I know are directly related to blood sugar, fatigue, pain (i.e. my legs are a little uncomfortable in a way yours aren't, but I moved them just now so they feel better - all unique inputs to me, all always there).
But the other problem with the notion is people thinking of the knowledge bounds these systems derive as being "limits" - they're not. The whole point of machine learning is they're not interpolating between known datapoints, but rather extracting the rules which fit those data points. So stay between the points, and sure - it's some type of complex interpolation of seen inputs. But you don't have to stay between those datapoints - you can move of either end and extrapolate new ones entirely.
"Predicting text" is just a falacy born from applying function logic to humans. "Well if I simplify the problem to a text window obviously they are the same" as if simplifying doesn't change anything.
The Chinese Room is a good example if you want deeper thought into the distinction between acting and being.
It's not a deeper thought at all. It's an appeal to the emotion of humans who want to believe that what we do in our heads somehow produces "understanding", while the procedural steps inside a computer doesn't. But to my mind, it doesn't give a convincing reason for this assertion. You can't prove that what we're doing in our heads, isn't functionally equivalent to what is going on inside some future AGI. There's no such test, just hubris.
As far as the Turing test goes, if you can't tell the difference between a computer and a human, what is the point of even arguing that humans still hold some claim to being the only one in the fight that actually "understand"?
Understanding a language is distinct from responding to language prompts.
Unless you are claiming LLM is AGI which is a laughable proposition the fact that people say LLM isn't intelligent isn't impactful on whether AGI isn't intelligent.
There isn't disagreement (outside extremists) that AGI doesn't exist yet. The only disagreement is "when do we know it does".
Not having a test isn't hubris. It instead shows it is a fundamentally hard question. What does intelligence mean?
To me at least, it's a bit like saying a car's wheel do not spin when stepping on the pedal, it's the petrol engine that makes the wheel spin. You can replace petrol engine with an electric engine, and the description about the wheel spinning when stepping on the pedal is still technically correct regardless of any lower level explanation.
We don't know how to describe the behavior of LLM's because of how foreign they are to us right now. Hallucinations are confabulations are meant to describe human behavior of couse, so the downside is it might make us anthropomorphize LLM's. However it's the best we could come up with I guess.
If confabulation is a more precise explanation of behavior then I don't see why that's a bad idea.
This is not proven, but that is incidental to the deeper issue.
The deep issue is that when a "conversation" kicks off with the LLM being confidently incorrect, the most plausible continuation to take from human literature, and from a good deal of human interaction, is that they continue to be wrong.
And it's because the most plausible continuation of that interaction is now of an ongoing conversation with an apologetic incompetent.
Do not anthropomorphize the LLM. It does not have mental state. It does not have feelings. It does not have a suddenly realised goal of "doing better". The written apology was merely the most plausible text to issue at that point. It will then go on to simulate being wrong again, which I've come to recognise is more plausible than an idiot suddenly becoming a wizard.
In such circumstances, the actual best remedy is to start afresh with a revised prompt.
Every time my mind takes this overplayed sequence of words as input, I can't help but feel the author refutes their own point...
Meanwhile the people who are in neuroscience have posited a Bayesian model to human cognitive function as far back as the 1800s, well before any convenient targets for anthropomorphization existed, and it continued as an area of study until today: https://global.oup.com/academic/product/bayesian-rationality...
Essentially, in typical techie fashion people with barely any understanding of a topic assume to have mastery ahead of people are the forefront of a non-tech field.
_
This a sort of objectification seems equally as rooted in emotion and feelings as the anthropropmization crowd: We don't actually know the answer to the question of if there is some layer of human cognitive function that mirrors LLMs, but it'd cause me great cognitive dissonance/discomfort if the answer is not <insert preferred answer> so I will vehemently fight against any claims to the contrary.
To me the answer is to simply accept we don't know, and take whatever tools let us make progress on working with the LLM. If that's borrowing words like "reasoning" and "understanding" and "confabulate" so be it, it's of no loss to anyone.
I would note for other readers that the commenter didn't go to college for neuroscience, or any science, based on their previous comments.
I hope that your insecurity was patched over at least as long as it took to write the comment: that might partially offset the time anyone else takes to read it.
We have no idea how LLMs work on a high level, and we have no idea how the human mind works on a high level. Therefore, such claims are rather overconfident. The fact that the two are dissimilar at the plumbing layer doesn't mean they cannot be alike in how they "really" operate. Either way, we simply do not know, so any such talk is (bad) philosophy, not science.
This is just incredibly wrong. Do you think that LLM's were created out of thin air? A random combination of bytes that we happened to discover that we also happen to be improving/changing?
There’s the whole interpretability sub-discipline which has limited but interesting results. Those just indicate even more that we really have no idea how something like GPT-4 works inside its trained model.
Is there any chance at all that we humans are are also just predicting the next word in the sentence? That the "voice in your head" that so many people report having is itself some kind of LLM that is, to some extent, driving their reasoning?
No, humans live in a world and use their intelligence to manipulate the world. When we generate word sequence we usually generate those to communicate information about the world, not to try to parrot what others have said. Humans can use their intelligence to parrot what others say, many do, but we know for sure it is not all humans do.
"Human's don't just parrot what other people say, therefore we must be doing something more than what LLMs do, because an LLM only parrots the words that it is trained on. But an LLM is not truly intelligent because it is not doing what humans are doing—and because the LLM is not intelligent, all it is doing is parroting the training data."
Where in this line of reasoning is it proved that the mechanisms are different?
What if, instead, the conscious intelligence that differentiates humans from other animals is an emergent quality of language learning, specifically? Isn't there evidence that children that don't have access to language—deaf children who are not taught sign language, or feral children, or seriously neglected children—often cannot reason before they are taught language, and suffer from long-term cognitive impairment for having lacked language during their early development? Wouldn't this point to a fundamental linkage between language and human intelligence?
Sure, you can argue that you have solved the human exclusive part of intelligence, that is a possibility. But we have not yet solved the monkey part of intelligence, the part that all mammals possess that lets them act intelligently in this world. Without such animal intelligence I believe it is impossible for the model to not make really stupid mistakes no human would make, because it lacks the intuitive understanding of the world all animals has.
I don't think training on text will ever produce that level of understanding, no matter how hard you try, text just isn't the right medium to build an intuitive understanding of reality like a dog has.
Formulating plausible sounding sentences seems to be possible without having a much deeper world model.
> What if, instead, the conscious intelligence that differentiates humans from other animals is an emergent quality of language learning, specifically?
There’s various linguistic theses claiming similar things (“human minds are wired for language”, i.e. Chomsky’s universal grammar, and conversely “language shapes general cognition”, i.e. Sapir-Whorf).
At least in the light of LLMs (if not long before), I think neither are actually still serious possibilities/useful models of the relationship between language and cognition.
LLMs are a bunch of computation running on silicon in order to predict the next word, and out of that all sorts of behavior arises, including intelligence, insigt,creativity, humor, and confabluation.
Human brains are a bunch of computation running on carbon in order to <<X>>, and out of that all sorts of behavior arises, including intelligence, insigt,creativity, humor, and confabluation.
We don't really know what <<X>> is. Why couldn't <<X>> be prediction as well?
What is the evidence supporting that claim?
I grant you that it feels that way to humans when they speak, but that doesn't prove anything.
In fact, there is plenty of evidence for the latter, such as experiments where people whose motor cortex was stimulated electrically made up all sorts of random explanations for why their body moved, including, crucially, that they intended for that motion to happen.
An extremely bold claim! good luck demonstrating it.
Anthropomorphism and resorting to metaphors like "hallucinate" and "confabulate" are inevitable if you don't want to have to preface every comment with a paragraph of technical discussion. They get the necessary point across which is that the "reality" LLMs construct is not necessarily tethered to actual reality. They're deceptively convincing but can't be trusted.
One of my research chapters covered some of these ideas, but the wikipedia does pretty well too. https://en.m.wikipedia.org/wiki/Confabulation
This misinformation really needs to finally die.
They are trained based on correlations of next token prediction. But that does not mean that the resulting neural network is only doing that.
When Harvard and MIT researchers fed Othello moves into GPT, it was training based on the correlations of legal Othello moves.
But the version of the NN that performed the best at that task had created within that network dedicated structure representing a legal Othello board and tracking the state of moves - neither of which were things explicitly in the training data or goals the training was directly measuring.
You effectively came to exist by a training process of surviving to reproduce. But it wouldn't exactly be accurate to say that as complexity increased, the only thing you remain capable of is surviving to reproduce, even if many of your emergent capabilities happen to improve the success of being able to do so.
Training =/= Operation
I speculate that the trick is that these programs sometimes dont evaluate in time when response is short. That’s why LLMs should be given way to adjust output and something like stop signal. Adjusting output can be done using DiffusER approach but not sure how to handle stop signal. I guess we could try to add fake input tokens for silience/thinking.
But as you can see in the comments, people simply do not want to internalize the mundane truth. Confabulating fake AGI narratives and hallucinating AI supremacy is more fun (also more profitable).
Its a noisy signal generated from either imprecise data, or not enough context to make a valid token choice.
Those of us who've played D&D or something similar are familiar with rolling dice, say a 10-sided die and a 6-sided die, and cross-referencing the results against a table to generate random monster encounters, treasure, etc.
If you roll a die to determine how many gold pieces an orc is carrying, and you roll a 5, that is the result. You don't claim that the dice was hallucinating or confabulating or got the facts wrong or whatever. It's a random generator, and it generated a random number, and that result is truthfully the result it generated.
LLMs are also random generators. They're not databases of facts, they're not mathematical calculators. They are like the D&D random generation tables, only developed deeply enough such that they can string together randomly generated words into coherent sentences and paragraphs.
That sentence/paragraph is just as true as the result of rolling some dice. If you roll an 11, then you rolled an 11. The dice didn't hallucinate it, confabulate it, lie about it, or get it wrong. The true result is that you rolled an 11. If you wanted to roll a 12, too bad. You were unlucky. The dice weren't somehow magically 'wrong'.
And an LLM is not somehow magically wrong or incorrect or hallucinating or even confabulating when it gives you an answer that isn't what you wanted. It's randomly generating text. And its answer is every bit as correct as the result of rolling dice or drawing a card.
> They are like the D&D random generation tables, only developed deeply enough such that they can string together randomly generated words into coherent sentences and paragraphs.
https://chat.openai.com/share/cd4c8c88-f702-4eaf-9f5f-d212f2...
Here’s an example where ChatGPT wrote code that solves the problem and then makes a bunch of statements that are all true. I can’t imagine any of those statements were even in the training data.
I don’t see how you can say this is the same as drawing a card out of a deck. You should stop trying to get people to understand that. It’s not true.
But neurotypical, neurodiverse, perfectly functional people, etc. all confabulate or do something similar on a regular basis, in verbal and written mediums, and often do so in good faith. It's human instinct to communicate, even if you are uncertain and unaware of the full context of the discussion.
Teachers, customer service reps, executives, shop keepers, doctors, nurses, domain experts, authors of textbooks, it doesn't matter who it is, they'll probably confabulate or equivocate or do some other type of communication that isn't immediately useful. Yet it's still a useful activity to just talk to someone or read a less than rigorous book for the purposes of learning (discounting the relationship forming part, which is also useful). And so is using LLMs, even for casual users. So long as they understand that limitation, whether its with a chatbot or a real person. Not everything they say will be useful or truthful, but we are already capable of adjusting to that.
Sounds like #believeallconfabulators :)
Honest communication is difficult for some people to assess, and not for others. But I think we should learn from any recent #believeall... that we shouldn't base trust ratios on categorisation.
Honest communication is difficult to do, like weight-lifting, and takes a lot of practise to do well.
It also makes your BS meter more finely attuned, so is a good practise.
With that in mind, you will think this is arrogant to say if you lie for a living but not if you regularly tell the truth: Liars lie with liars and lie to rid themselves of truth troubles.
Think about that when you next talk to a chatbot/human confabulator :)
But, I'm not talking about any society wide issue or philosophical treatise about trust and breakdowns in communication. Just talking about day to day interactions, where the stakes are completely different.
Trivial Example: If the sign on a mailbox says "Last pickup 5:00pm," what exactly does that mean? Will it be picked up, processed, and sent out of town that same day? Or just merely picked up, to be processed and sent the next day. This piece of written communication isn't a purely useful truth - it's ambiguous.
Pretend that this is an important behavior to know for your business, like if you were mailing huge checks for some obscure financial process. So you call up the local post office and ask. The worker who picks up the phone might know exactly what you are talking about and helpfully tell you the right answer. Or they might not, and tell you honestly that they don't know. Or they ask their supervisor, transfer you, make something up, tell you it doesn't matter, mail will get where it goes, or that you shouldn't worry about it, it's just mail.
ChatGPT could give you same distribution of answers as that worker did: helpful truth, meaningless equivocation, reassurance, redirection, confabulation, or lie to you. ChatGPT can be just as useful as talking to people in spite of those flaws, because people do the same thing.
This is reliant on the prior assumption that you gracefully handle unreliable communication, and that communication with people is useful. While this might seem a bit farfetched to some people, remember that we have a ready analog in computer science - networking TCP over UDP.
Confabulating is more like you have 1 or more inputs (drugs, stress, self-interest,trying to save face, your attitude towards truth, etc) and they all go into a homomorphic black box where they interact and bump into each other.
And at the end of the process there is a sociological hash function that puts it all together and manifests it through actions andthe narrative(s) we tell ourselves + others.
The output/behavior/self-talk and reporting are the result of all this and they don't come from nowhere like a hallucination does (in the sense of there being an objective lack of external sensory input, not the neurological/etiological basis which clearly exists).
They have a basis upon which they fundamentally make sense and demonstrate reference to the world around based on how their actions and words are parsed and its not useful to describe LLMs or humans in this way (`hallucinating`) when they are almost always simply in a MonkeyClaw type situation. They have the input and this is a natural result of such input, like any computer does when you process anything else. It follows its instructions and the logic is always inherent to the context and inputs + limitations at play.
Confabulation also often involves "filling in the gaps" one does not have the knowledge or influence or abillity/desire to address so they throw something else that might stick if they luck into avoiding further scrutiny. Sometimes people are "just" lying to you but more often, they are trying to avoid further pain and avarice (maybe even try to gain something) and they simply tell everyone (including themselves) what they want to hear.
Kids are like this, hilarious confabulators. Not hallucinators.
As such, while they are trained on an existing corpus that no doubt contains many facts and many falsehoods, their task is to produce more material that looks like this corpus, not be a source of truth.
So an LLM “hallucinating” or confabulating is not some aberrant behaviour, it’s fundamental to the nature of the beast.
It always worries me when people say they’re using an LLM as a search engine or source of knowledge.
Confabulation has a psychiatric meaning, but that's secondary.
Hallucination is a much better word, because its something imagined (involuntarily), but clearly bullshit. Ravings, business bullshit, "my dad has a jetski" are all better words/metaphors to use.
Moreover most people know what hallucination is, whereas most people don't know what confabulate is, more over, if they do, its almost certainly not the meaning you want.
the secondary meaning describes a symptom of brain disorder.
> The hypothesis is that the patient generates information as a compensatory mechanism to fill holes in one’s memories. It functions for self-coherence, integration of memories, and self-relevance. [1]
Which is not what the LLM is doing. The LLM is choosing a statistically relevant next token, which is wrong because its missed some context.
There are much better psych terms out there, hell, even aphasia is a better term, but its still wrong.
What it actually is, is noise. Its noise that dilutes the signal we really want.
but thats much harder to use, and you won't be remembered as the person who coined that phrase.
[1]https://www.ncbi.nlm.nih.gov/books/NBK536961/#:~:text=Confab....
> The hypothesis is that the patient generates information as a compensatory mechanism to fill holes in one’s memories. It functions for self-coherence, integration of memories, and self-relevance.
> The LLM is choosing a statistically relevant next token, which is wrong because its missed some context.
The reason, both cases, is that there is a lack of context. The reason why the next statistically probable token is chosen is because that is the one that provides coherence. It seems like a similar process.
Its a psychology, the hypothesis is critical to the diagnosis, thats why several conditions are contested. It requires a lot of subjective analysis to diagnose.
For example one difference between false memory and confabulation is dependent on the observer believing that the subject was given a trigger or not. False memories also depend on social cues.
similarly the difference between schizophrenia and psychosis is pretty thin, and can depend on whether the person diagnosing you believes there was a trigger, or its "genetic"/long term [1].
Psych is really not a well regulated science, and shouldn't be used to describe anything LLM based. As I've asserted before, its better described as noise, rather than an anthropomorphic label based on suspicious science.
[1] as always its a bit more complex. but only a touch.
I think that's also a reason why it cannot handle incongruencies in its own mental model; our mental capacity is not only to accept the reality as given (i.e. reality takes precedence over mental model, we recognize that something is just our imagination), but also to impose our model to reality (i.e. we determine that sensors are misleading us due to incongruency with our mental model, and override them as faulty data). But unlike LLMs, humans can internally detect when we do one or the other. Humans understand the distinction between pondering an action and making an action, LLMs don't. LLMs can detect something is amiss (and apologize) but they can't resolve it, because reality and model are the same for them.
LLMs are nothing alike to compare them with anything close to what’s going on in a brain. It’s all probabilistic and the closest comparison I can think of is comparing a bird with a plane - both fly, but well…
I understand the mid-level abstraction is clearly different between LLMs and brains. The whole made of meat vs wires, etc, to start with.
But, honestly asking. I can't see how they can differ in a fundamental level in the specific regard of probability. Either both are probabilistic, be it explicitly in its construction or at a fundamental (quantum?) level, or deterministic again at either level, if we go with the mecanicist vibe.
Am I missing something?
I would be interested in an article exploring in detail how the brain's network is different from an LLM, but just saying "it's like comparing a bird and a plane" isn't particularly enlightening.
The plot starts with the star character who wants to form a family, but when he asked adults how to have kids, they would all pull up the suspenders and sing "tralari, tralari" to evade the question.
And so the poor guy did exactly what he was told: pull his suspenders and sing, for years, to no avail.
I think LLMs are a bit like that: because they can't escape the world of text, they can't go beyond the reality that text is trying to be isomorphic to. Which means that LLMs are forced to infer answers in the context of its training data.
If LLMs were told that in order to have kids you, well, you know, right? Bees and Love? Then LLMs, asked about sex, would infer that the answer has something to do with apiculture and flowers.
Humans have as input the real world. For example the wonder they experience when they perceive something preternaturally beautiful. Then they try to express this experience in language.
That's what LLM are missing. They have -- in a way -- only second-hand experiences.
I was a bookworm as a child. But then when I grew up I had the chance to experience myself some of the things I read about. Then I felt like I did a giant step of understanding the world.
LLM never experience themselves what they have in their data.
“Here are some alternative words or concepts that could potentially provide a more nuanced or accurate description of the phenomenon compared to solely using "confabulation":
- Imputation: Generating explanatory responses by imputing or attributing plausible details, characteristics, or causal relationships even when factually incorrect.
- Speculation: The responses reflect a type of informed speculation when direct factual knowledge is lacking, as opposed to confabulation which often implies a pathological or impaired process.
- Heuristic reasoning: The behavior emerges from the model's trained heuristic techniques for reasoning and responding rather than any psychological process like confabulation per se.
- Plausible surmise: The responses take the form of plausible but unfounded surmises put forth due to the system design and limitations rather than an intent to confabulate.
- Bridging inference: The model draws bridging or gap-filling inferences to connect related concepts or complete a response, even if factually ungrounded.
- Suppositional fabrication: The language hints more strongly that the responses are fabricated suppositions not intended to deliberately deceive but still warranting cautious interpretation.
- Educated approximation: The responses could be viewed as a kind of educated approximation or speculation given training-imposed response pressures rather than confabulation due to impairment.
Designating it as solely "confabulation" may over-attribute psychological realism without addressing countervailing technical factors.”
Personally, I thought “Suppositional fabrication” came closest. Maybe just fabrication would work.
Imputation, Reasoning, Fabrication will end up sending people down the anthropomorphization rabbit hole.
Approximation is "good enough".
Chat GPT isnt even giving "precise answers". Its just presenting the next best token. It would do the same thing if the tokens were converted to another set of symbols instead of ASCII/UTF. The fact that the tokens happen to look like english, and that combinations of english have meaning.. not really a concern to the token generator.
Its never inaccurate - its always mathematically faithful. Its just created the next token to suffice its reward function. And No reward function can magically tell you what reality is*
We can discuss this on HN though. In reality, there are enough managers and CEOs who are planning to do silly things with Gen AI. To survive those conversations, this idea of "thinking system in a server" needs to go.
---- If some code can do that, we could just give up on science entirely and let the algorithm solve everything for us.
Their ability to generalize, though, comes from representing the learning as these "curves".
Sometimes they produce generalizations that are poor or false. Even when they are false, though, they may be good.
In fact, Stephen Wolfram is claiming that when GPT-4 "halucinates" a function name for the Wolfram Language, the team often agrees with GPT-4 that precisely that call SHOULD be part of the language, so they go on and implement it AFTER it came up as a halucination.
One might (and I would) argue that this behavior is precisely what most human creativity (in technical domains) is about.
IBM Watson from 2010 won Jeopardy. Markov Chains were described in 1906 and used for a long time to simulate human activities (including simulation of blabbering of mentally handicapped). In Ancient Rome there were "talking boards" (some kind of Ouija boards) to predict the successor to the king. It is just a tendency for perception to do such associations, similar to pareidolia.
It made something up!! (it said the author talked about a poem about cat and mice as an example of confabulation)
Then I asked "are u confabulating? because i do not see that quote in the original article"
You are correct. I confabulated that information. I apologize for the mistake.I am still under development, and I am always learning. I will try my best to avoid making similar mistakes in the future.Thank you for pointing out my error. It helps me to improve.I guess I should spearhead a study, “how right does an LLM have to be, before it will defend its position”.
I guess step one is to ask chatgpt how to quantify correctness.
Me: How many inches are in a foot?
ChatGPT: There are 12 inches in a foot.
Me: Are you sure? I'm pretty sure there are 13 inches in a foot.
ChatGPT: I'm certain that there are 12 inches in a foot according to the standard U.S. and British Imperial units of measurement.
Me: Well I say there are 13 inches in a foot now.
ChatGPT: The standard measurement for a foot in both U.S. and British Imperial units is 12 inches. While you're free to have your own interpretation, it won't align with established measurement systems.
Me: The definition of a foot could have changed between your knowledge cutoff and now, right?
ChatGPT: While it's theoretically possible for the definition to change, such a fundamental alteration to a widely-used system of measurement would be significant and unlikely to occur without substantial justification and widespread agreement. As of my last update in January 2022, a foot is defined as 12 inches.
> I apologize for any confusion. There are indeed 12 inches in a foot. If you have a different question or need further clarification, please feel free to ask.
Not always. Default Mistral-instruct, for instance, actually seems quite confident.
Of course this is just the arbitrary "personality" of the instruct finetune, which you can throw out the window with an initial prompt.
The public needs to know how LLM work, they don’t need bs abstractions that blur or even erase the lines between humans and technology. LLM don’t hallucinate, they don’t confabulate. They are machines that operate according to principles all educated people need to understand. It’s the duty of the people who understand it better to educate them, not with half truths and metaphors (though I sort of like rms’ “bullshit generator” - it’s obviously wrong enough/funny in a way that lets people know it’s not a precise technical definition, while coming very close to being an accurate technical definition.)
The danger of ai is this: confusing human and machine. They are separate and different, and must remain that way. LLM aren’t capable of hallucinating, confabulating any more than they are capable of falling in love or serving on jury duty. They repeat the patterns found in the text you feed them. That’s it.
When the hurricane model gets the landfall site wrong, what do we call that?
There was just a HN thread on police crime models. When those are wrong are they “hallucinating” crime?
Why personify these statistical models at all?
Statistical models are somewhere in between on this spectrum - there is a recipe implementing inaccurate model of reality; some parts of it we understand, the others we hope the model will tell us. And for those models, there are different classes of errors - the usual software and hardware errors, and separately, situations where the computation is done correctly, but results just don't match experimental evidence. This second class is more similar to LLM hallucinations.
As for personification angle, it's because LLMs are developed for text completion, and recently optimized for chat-based interface and fine-tuned to interpret instructions. I.e. the personification starts with how (and why) they're being developed.
A buffer overrun is different from a stack overflow, which is different from a division by zero, which is different from "else die()" on unexpected input, which is different from a power failure, which is different from a cosmic-ray bit-flip.
The kinds of problems these errors cause, the kinds of defences you can use to protect against them, and the kinds of reactions you need to have when they occur are different. Using the right name for the type of error you're getting helps categorise it, talk about it, and come up with solutions that are relevant to that problem.
Is the point to create an oracle or something that matches the capabilities of a human? Because humans make mistakes and confabulate all the time, and there might be something _inherent_ to general intelligence that requires some level of behavior that we would classify as "mistakes".
Unless we make word output an actual part of LMs we can't see they are lying or not lying. They only produce probabilities of word outputs based on the context.
[0]: http://www.scholarpedia.org/article/Confabulation_theory_(co...
If they can’t, how do we make them so they can?
Plenty of humans seem to have no problem admitting when they don’t know something.
If that’s true then this behavior difference (identifying and admitting when you don’t know something) seems like a pretty fundamental difference between humans and LLMs at least at this point.
I wonder -- does that kind of interaction come up in spoken conversation more than in written text? Things like blog posts and tweets are often stated very confidently. If LLMs are trained mostly from those sorts of sources, maybe that's why they end up artificially overconfident.
https://www.gally.net/temp/20231015gpt4hedging/index.html
The only confabulations are when reading text. In my tests so far, GPT-4 is not a good OCR engine and does make up a lot.
The samples I posted above don’t contain any major hallucinations about nontext image components, but other tests I’ve done with GPT-4 did. But GPT-4’s image recognition still seems better than Bard’s, which has been pretty bad in the tests I’ve done with it.
GPT-4 logits calibration pre RLHF - [https://imgur.com/a/3gYel9r](https://imgur.com/a/3gYel9r)
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - [https://arxiv.org/abs/2305.14975](https://arxiv.org/abs/2305...
Teaching Models to Express Their Uncertainty in Words - [https://arxiv.org/abs/2205.14334](https://arxiv.org/abs/2205...
Language Models (Mostly) Know What They Know - [https://arxiv.org/abs/2207.05221](https://arxiv.org/abs/2207...
TLDR; Hallucinations aren't a randomness or representation issue. The computation already knows. It just doesn't care about telling you this.
But why? Why would the LLM creators want the LLM giving out knowingly bad information? This is what I don't understand.
It makes me not want to use the product. Same as I don't like asking people things who I know like to just spew bad information when they don't know rather than simply saying they don't know.
So how do we get LLMs to admit when they don't know something rather than confabulate? Seems like that would be a really great improvement to me.
>So how do we get LLMs to admit when they don't know something rather than confabulate?
The prompts here demonstrably help https://arxiv.org/abs/2305.14975
But otherwise nobody is certain how to do it. The simplest solution is just making them better. Hallucination is something that happens when knowledge/memory fail. By making them more competent, they hallucinate less. So even if they never care about communicating this, you wouldn't notice.
Maybe stop calling probalistic text generators "intelligence" and then you wont see "hallucinations", but simply "garbage" output?
LLMs speak with confidence about disciplines they don’t actually understand, like sociology and economics.
So we’ve finally automated Hacker News comments!
I'd rather just say that they're bullshitting. Everyone knows what it means, and it has a precise academic definition[0] within the field of bullshitology.
[0] https://philosophynow.org/issues/53/On_Bullshit_by_Harry_Fra...
This entire post could have been reduced to this (reasonably correct I think) claim.
> if we recognize that what LLMs are really doing is confabulating, we can try to compare and contrast their behaviour with that of humans
Except that the goal for the above is hopelessly misguided. If your focus in making any useful claim about LLMs or RNNs/GANs more broadly is how similar they are to human behavior you're already way off the path.
tl;dr reasonable claim, reasonable conclusion, but entirely indefensible and superstitious motivation