Detecting hallucinations in large language models using semantic entropy
nature.com
nature.com
For any given input text, there is a corresponding output text distribution (e.g. the probabilities of all words in a sequence which the model draws samples from).
The approach of drawing several samples and evaluating the entropy and/or disagreement between those draws is that it relies on already knowing the properties of the output distribution. It may be legitimate that one distribution is much more uniformly random than another, which has high certainty. Its not clear to me that they have demonstrated the underlying assumption.
Take for example celebrity info, "What is Tom Cruise known for?". The phrases "movie star", "katie holmes", "topgun", and "scientology" are all quite different in terms of their location in the word vector space, and would result in low semantic similarity, but are all accurate outputs.
On the other hand, "What is Taylor Swift known for?" the answers "standup comedy", "comedian", and "comedy actress" are semantically similar but represent hallucinations. Without knowing the distribution characteristics (e.g multivariate moments and estimates) we couldn't say for certain these are correct merely by their proximity in vector space.
As some have pointed out in this thread, knowing the correct distribution of word sequences for a given input sequence is the very job the LLM is solving, so there is no way of evaluating the output distribution to determine its correctness.
There are actual statistical models to evaluate the amount of uncertainty in output from ANNs (albeit a bit limited), but they are probably not feasible at the scale of LLMs. Perhaps a layer or two could be used to create a partial estimate of uncertainty (e.g. final 2 layers), but this would be a severe truncation of overall network uncertainty.
Another reason I mention this is most hallucinations I encounter are very plausible and often close to the right thing (swapping a variable name, confabulating a config key), which appear very convincing and "in sample", but are actually incorrect.
It depends on the fact that a high uncertainty answer by definition is less probable. That means if you ask multiple times you will not get the same unlikely answer, such as that Taylor swift is a comedian, you will instead get several semantically different answers.
Maybe you’re saying the same thing, but if so I’m missing the problem. If your training data tells you that Taylor Swift is known as a comedian, then hallucinations are not your problem.
A better example might be that the model overtrained on AWS cloud formation API 2 and when v3 comes out produces low entropy answers that are wrong for v3 but right for v2 (due to training bias), but the answers are low variance (e.g. "bucket" instead of the new "bucket_name" key).
Another example based on a quick test I did on GPT4:
In a single phrase, what is Paris?
Paris is the city of Light.
Paris is the capital of France.
Paris is the romantic capital of the world renowned for its art, fashion, and culture.
I don’t know the technical term for it, if there is one, but essentially by prompting with “In a single phrase” you are over-constraining the answer space, and so again you’ve moved into a high uncertainty regime.
Consider if you prompted a tool to render Paris in a single image. Eiffel Tower and Arc? Paris street with bakery? Overhead maps view? With what included and what omitted?
Hallucinations are an error that occurs based on providing an answer with high uncertainty, not wrong or outdated training data.
Blaming human semantics for LLMs is generally a bad idea, since we only use human semantics to qualitatively explain how the models abstract ideas. In practice you simply don't know how the model relates words.
To me this sounds very similar to lowering temperature. It doesn't sound like it pulls better from grounded-truth but rather more probabilistic in the vector space. Does this jive?
For your Tom Cruise example, since all those phrases are true and grounded in the training data, the technique may fire off a false positive "hallucination decision".
However, the example they give in the paper seems to be for "single-answer" questions, e.g., "What is the receptor that this very specific medication acts on?", or "Where is the Eiffel Tower located?", in which case I think this approach could be helpful. So perhaps this technique is best-suited for those single-answer applications.
> draw[ing] several samples and evaluating the entropy and/or disagreement between those draws
the method from the paper (as I understand it):
- samples multiple answers, (e.g. "music:0.8, musician:0.9, concert:0.7, actress:0.5, superbowl:0.6")
- groups them by semantic similarity and gives them an id ([music, musician, concert] -> MUSIC, [actress] -> ACTING, [superbowl] -> SPORTS), note that they just use an integer or something for the id
- sums the probability of those grouped answers and normalizes: (MUSIC:2.4, ACTING:0.5, SPORTS:0.6 -> MUSIC:0.686, SPORTS:0.171, ACTING:0.143)
They also go to pains in the paper to clearly define what they are trying to prevent, which is confabulations.
> We focus on a subset of hallucinations which we call ‘confabulations’ for which LLMs fluently make claims that are both wrong and arbitrary—by which we mean that the answer is sensitive to irrelevant details such as random seed.
Common misconceptions will still be strongly represented in the dataset. What this method does is it penalizes semantically isolated answers (answers dissimilar to other possible answers) with mediocre likelihood.
Now technically, this paper only compares the effectiveness of "detecting" the confabulation to other methods - it doesn't offer an improved sampling method which utilizes that detection. And of course, if it were used as part of a generation technique it is subject to the extreme penalty of 10xing the number of model generations required.
link to the code: https://github.com/jlko/semantic_uncertainty
> If it simply output "Capital of Rome" 3 times, would that indicate its probably not a confabulation? Correct, it would tend to indicate that. Whether or not that is a true output is a different question, but it would tend to indicate that the "Capital of Rome" is something that comes up consistently regardless of variations in random seed.
It's not a solution for hallucinations writ large, and it doesn't introduce an oracle of ground-truth accuracy on factoids. It just reweights possible answers by their 'semantic entropy'.
They did something else cool in the paper too (unrelated to whatever problem you want to solve with respect to apriori knowledge) They made a method for decomposing originally generated answers into "factoids" and fact checking them using this semantic entropy concept:
- Generate an output to a question e.g. "who is tom cruise?"
- for each sentence fragment that represents a factoid "known for the movie top gun", make a question jeopardy style: "what fighter pilot movie is tom cruise known for?"
- for each question, generate multiple answers and select one with low semantic entropy
- compare that answer with the original factoid
pretty cool overall, but still doesn't get us any closer to LLMs with a sense for whether something is fact-ish / a sense of truth.
Taylor Swift has appeared multiple times on SNL, both as a host and as a surprise guest, beyond being a musical performer[0]. Generally, your point is correct, but she has appeared on the most famous American television show for sketch comedy, making jokes. One can argue whether she was funny or not in her appearances, but she has performed as a comedian, per se.
Though she hasn't done a full-on comedy show, she has appeared in comedies in many credits (often as herself).[1] For example she appeared as "Elaine" in a single episode of The New Girl [2x25, "Elaine's Big Day," 2013][2]. She also appeared as Liz Meekins in "Amsterdam" [2022], a black comedy, during which her character is murdered.[3]
It'd be interesting if there's such a thing as a negatory hallucination, or, more correctly, an amnesia — the erasure of truth that the AI (for whatever reason) would ignore or discount.
[0] https://www.billboard.com/lists/taylor-swift-saturday-night-...
[1] https://www.imdb.com/name/nm2357847/
[2] https://newgirl.fandom.com/wiki/Elaine
[3] https://www.imdb.com/title/tt10304142/?ref_=nm_flmg_t_7_act
The process can look like-
1. Use existing large models to convert the same previous dataset they were trained on into formal logical relationships. Let them generate multiple solutions
2. Take this enriched dataset and train a new LLM which not only outputs next token but also a the formal relationships between previous knowledge and the new generated text
3. Network can optimize weights until the generated formal code get high accuracy on proof checker along with the token generation accuracy function
In my own mind I feel language is secondary - it's not the base of my intelligence. Base seems more like a dreamy simulation where things are consistent with each other and language is just what i use to describe it.
How to extend LLMs to add mechanisms for reasoning, causality, etc (Type 2 thinking)? However that will eventually be done, the implementation must continue to be probabilistic, informal, and bottom-up. Manual human curation of logical and semantic relations into knowledge models has proven itself _not_ to be sufficiently scalable or anti-brittle to do what's needed.
We could just use RAG to create a new dataset. Take each known concept or named entity, search it inside the training set (1), search it on the web (2), generate it with a bunch of models in closed book mode (3).
Now you got three sets of text, put all of them in a prompt and ask for a wikipedia style article. If the topic is controversial, note the controversy and distribution of opinions. If it is settled, notice that too.
By contrasting web search with closed-book materials we can detect biases in the model and lacking knowledge or skills. If they don't appear in the training set you know what is needed in the next iteration. This approach combines self testing with topic focused research to integrate information sitting across many sources.
I think of this approach as "machine study" where AI models interact with the text corpus to synthesize new examples, doing a kind of "review paper" or "wiki" reporting. This can be scaled for billions of articles, making a 1000x larger AI wikipedia.
Interacting with search engines is just one way to create data with LLMs. Interacting with code execution and humans are two more ways. Just human-AI interaction alone generates over one billion sessions per month, where LLM outputs meet with implicit human feedback. Now that most organic sources of text have been used, the LLMs will learn from feedback, task outcomes and corpus study.
My take - Hallucinations can never be made to perfect zero but they can be reduced to a point where these systems in 99.99% will be hallucinating less than humans and more often than not their divergences will turn out to be creative thought experiments (which I term as healthy imagination). If it hallucinates less than a top human do - I say we win :)
(The term "weed" for marijuana is just a joke derived from that sense of the word.)
[1] https://cyc.com
I was not actually aware that building KG from text is NP-hard problem. I will check it out. I thought it was a time consuming problem when done manually without LLMs but didn't thought it was THAT hard. Hence I was trying to introduce LLM into the flow. Thanks, will read about all this more!
Eg, KGs (RDF, PGs, ...) are logical, but in automated construction, are not semantic in the sense of the ground domain of NLP, and in manual construction, tiny ontology. Conversely, fancy powerful logics like modal ones are even less semantic in NLP domains. Code is more expressive, but brings its own issues.
LLMs generate responses based on statistical probabilities derived from their training data. They do not inherently understand or store an "absolute source of truth." Thus, any KG bootstrapped from an LLM might inherit not only the model's insights but also its inaccuracies and biases (hallucinations). You need to understand that these hallucinations are not errors of logic but they are artifacts of the model's training on vast, diverse datasets and reflect the statistical patterns in that data.
Maybe you could build retrieval model but not generative model.
I don't have answer to that. I felt there would be lesser KG representations which would fit a logical world, than what fits into the current vast vector spaces of network's weight and biases. But that's just a idea. This whole thing stems from this internal intuition that language is secondary to my thought process and internally I feel I can just play around concepts without language - what kind of Large X models will meet that kind of capability I don't know!
Why should you believe the output of the LLM just because it is formatted a certain way (i.e. "formal logical relationships")?
One expression of that idea is in this paper: https://link.springer.com/article/10.1007/s10676-024-09775-5
LLMs are incapable of presenting things as truth.
The social media gurus don't help with these issues by claiming that non-intentional objects are going to cause humanity's demise when there are much more pertinent issues to be concerned about like global warming, corporate malfeasance, and the general plundering of the biosphere. Algorithms that lie are not even in the top 100 list of things that people should be concerned about.
How do you know whether something has “intentions”? How can you know that humans have them but computer programs (including LLMs) don’t or can’t?
If one is a materialist/physicalist, one has to say that human intentions (assuming one agrees they exist, contra eliminativism) have to be reducible to or emergent from physical processes in the brain. If intentions can be reducible to/emergent from physical processes in the brain, why can’t they also be reducible to/emergent from a computer program, which is also ultimately a physical process (calculations on a CPU/GPU/etc)?
What if one is a non-materialist/non-physicalist? I don’t think that makes the question any easier to answer. For example, a substance dualist will insist that intentionality is inherently immaterial, and hence requires an immaterial soul. And yet, if one believes that, one has to say those immaterial souls somehow get attached to material human brains - why couldn’t one then be attached to an LLM (or the physical hardware it executes on), hence giving it the same intentionality that humans have?
I think this is one of those questions where if someone thinks the answer is obvious, that’s a sign they likely know far less about the topic than they think they do.
No, I never assumed “all physical processes are computational”. I never said that in my comment and nothing I said in my comment relies on such an assumption.
What I’m claiming is (1) we lack consensus on what “intentionality” is (2) we lack consensus on how we can determine whether something has it. Neither claim depends on any assumptions about “physical processes are computational”
If one assumes materialism/physicalism - and I personally don’t, but given most people do, I’ll assume it for the sake of the argument - intentionality must ultimately be physical. But I never said it must ultimately be computational. Computers are also (assuming physicalism) ultimately physical, so if both human brains and computers are ultimately physical, if the former have (ultimately physical) intentionality - why can’t the latter? That argument hinges on the idea both brains and computers are ultimately physical, not on any claim that the physical is computational.
Suppose, hypothetically, that intentionality while ultimately physical, involves some extra-special quantum mechanical process - as suggested by Penrose and Hameroff’s extremely controversial and speculative “orchestrated objective reduction” theory [0]. Well, in that case, a program/LLM running on a classical computer couldn’t have intentionality, but maybe one running on a quantum computer could, depending on exactly how this “extra-special quantum mechanical process” works. Maybe, a standard quantum computer would lack the “extra-special” part, but one could design a special kind of quantum computer that did have it.
But, my point is, we don’t actually know whether that theory is true or false. I think the majority of expert opinion in relevant disciplines doubts it is true, but nobody claims to be able to disprove it. In its current form, it is too vague to be disproven.
[0] https://en.m.wikipedia.org/wiki/Orchestrated_objective_reduc...
You seem to be saying that because we have no clear cut way of determining whether people have intentions then that means, by physical reductionism, algorithms could also have intentions. The limiting case of this kind of semantic hair splitting is that I can say this about anything. There is no way to determine if something is dead or alive, there is no definition that works in all cases and no test to determine whether something is truly dead or alive so it must be the case that algorithms might or might not be alive but because we can't tell then me might as well assume there will be a way to make algorithms that are alive.
It's possible to reach any nonsensical conclusion using your logic because I can always ask for a more stringent definition and a way to test whether some object or attribute satisfies all the requirements.
I don't know anything about theories of consciousness but that's another example of something which does not have an algorithmic implementation unless one uses circular logic and assumes that the brain is a computer and consciousness is just software.
What is an "intention"? Do we all agree on what it even is?
> What can be implemented with computers and digital circuits are deterministic signal processors which always produce consistent outputs for indistinguishable inputs.
We don't actually know whether humans are ultimately deterministic or not. It is exceedingly difficult, even impossible, to distinguish the apparent indeterminism of a sufficiently complex/chaotic deterministic system, from genuinely irreducible indeterminism. It is often assumed that classical systems have merely apparent indeterminism (pseudorandomness) whereas quantum systems have genuine indeterminism (true randomness), but we don't actually know that for sure – if many-worlds or hidden variables are true, then quantum indeterminism is ultimately deterministic too. Orchestrated objective reduction (OOR) assumes that QM is ultimately indeterministic, and there is some neuronal mechanism (microtubules are commonly suggested) which permits this quantum indeterminism to influence the operations of the brain.
However, if you provide your computer with a quantum noise input, then whether the results of computations relying on that noise input are deterministic depends on whether quantum randomness itself is deterministic. So, if OOR is correct in claiming that QM is ultimately indeterministic, and quantum indeterminism plays an important role in human intentionality, why couldn't an LLM sampled using a quantum random number generator also have that same intentionality?
> You seem to be saying that because we have no clear cut way of determining whether people have intentions then that means, by physical reductionism, algorithms could also have intentions.
Personally, I'm a subjective idealist, who believes that intentionality is an irreducible aspect of reality. So no, I don't believe in physical reductionism, nor do I believe that algorithms can have intentions by way of physical reductionism.
However, while I personally believe that subjective idealism is true, it is an extremely controversial philosophical position, which the clear majority of people reject (at least in the contemporary West) – so I can't claim "we know" it is true. Which is my whole point – we, collectively speaking, don't know much at all about intentionality, because we lack the consensus on what it is and what determines whether it is present.
> The limiting case of this kind of semantic hair splitting is that I can say this about anything. There is no way to determine if something is dead or alive, there is no definition that works in all cases and no test to determine whether something is truly dead or alive so it must be the case that algorithms might or might not be alive.
We have a reasonably clear consensus that animals and plants are alive, whereas ore deposits are not. (Although ore deposits, at least on Earth, may contain microscopic life–but the question is whether the ore deposit in itself is alive, as opposed being the home of lifeforms which are distinct from it.) However, there is genuine debate among biologists about whether viruses and prions should be classified as alive, not alive, or in some intermediate category. And more speculatively, there is also semantic debate about whether ecosystems are alive (as a kind of superorganism which is a living being beyond the mere sum of the individual life of each of its members) and also about whether artificial life is possible (and if so, how to determine whether any putative case of artificial life actually is alive or not). So, I think alive-vs-dead is actually rather similar to the question of intentionality – most people agree humans and at least some animals have intentionality, most people would agree that ore deposits don't, but other questions are much more disputed (e.g. could AIs have intentionality? do plants have intentionality?)
I don't follow. If intentionality is an irreducible aspect of reality then algorithms as part of reality must also have it as realizable objects with their own irreducible aspects.
I don't think algorithms can have intentionality because algorithms are arithmetic operations implemented on digital computers and arithmetic operations, no matter how they are stacked, do not have intentions. It's a category error to attribute intentions to algorithms because if an algorithm has intentions then so must numbers and arithmetic operations of numbers. As compositions of elementary operations there must be some element in the composite with intentionality or the claim is that it is an emergent property in which case it becomes another unfounded belief in some magical quality of computers and I don't think computers have any magical qualities other than domains for digital circuits and numeric computation.
I don't see how that makes it a category error? Like, assuming that numbers and arithmetic operations of numbers don't have intentions, and assuming that algorithms having intentions would imply that numbers and arithmetic operations have them, afaict, we would only get the conclusion "algorithms do not have intentions", not "attributing intentions to algorithms is a category error".
Suppose we replace "numbers" with "atoms" and "computers" with "chemicals" in what you said.
This yields "As compositions of [atoms] there must be some [element (in the sense of part, not necessarily in the sense of an element of the periodic table)] in the composite with intentionality or the claim is that it is an emergent property in which case it becomes another unfounded belief in some magical quality of [chemicals] and I don't think [chemicals] have any magical qualities other than [...]." .
What about this substitution changes the validity of the argument? Is it because you do think that atoms or chemicals have "magical qualities" ? I don't think this is what you mean, or at least, you probably wouldn't call the properties in question "magical". (Though maybe you also disagree that people are comprised of atoms (That's not a jab. I would probably agree with that.)) So, let's try the original statement, but without "magical".
"As compositions of elementary operations there must be some element in the composite with intentionality or the claim is that it is an emergent property in which case it becomes another unfounded belief in some [suitable-for-emergent-intentionality] quality of computers and I don't think computers have any [suitable-for-emergent-intentionality] qualities [(though they do have properties for allowing computations)]."
If you believe that humans are comprised of atoms, and that atoms lack intentionality, and that humans have intentionality, presumably you believe that atoms have [suitable-for-emergent-intentionality] qualities.
One thing I think is relevant here, is "we have nothing showing us that there exist [x]" and "it cannot be that there exists [x]" .
Even if we have nothing to demonstrate to us that numbers-and-operations-on-them have the suitable-for-emergent-intentionality qualities, that doesn't demonstrate that they don't.
That doesn't mean we should believe that they do. If you have strong priors that they don't, that seems fine. But I don't think you've really given much of a reason that others should be convinced that they don't?
Computers have a formal theory and to say that a computer has intentions and can think would be equivalent to supplying a constructive proof (program) demonstrating conformance to a specification for thought and intention. These don't exist so from a constructive perspective it is valid to say that all claims of computers and software having intentions and thoughts are simply magical, confused, non-constructive, and ill-typed beliefs.
That's not true. To give a trivial example, a set or sequence of numbers is composed of numbers but is not itself a number. 2 is a number, but {2,3,4} is not a number.
> Computers have a formal theory
They don't. Yes, there is a formal theory mathematicians and theoretical computer scientists have developed to model how computers work. However, that formal theory is strictly speaking false for real world computers – at best we can say it is approximately true for them.
Standard theoretical models of computation assume a closed system, determinism, and infinite time and space. Real world computers are an open system, are capable of indeterminism, and have strictly sub-infinite time and space. A theoretical computer and a real world computer are very different things – at best we can say that results from the former can sometimes be applied to the latter.
There are theoretical models of computation that incorporate nondeterminism. However, I'd question whether the specific type of nondeterminism found in such models, is actually the same type of nondeterminism that real world computers have or can have.
Even if you are right that a theoretical computer science computer can't have intentionality, you haven't demonstrated a real world computer can't have intentionality, because they are different things. You'd need to demonstrate that none of the real differences between the two could possibly grant one the intentionality the other lacks.
That's still a number because everything in a digital computer is a number or an operation on a number. Sets are often encoded by binary bit strings and boolean operations on bitstrings then have a corresponding denotation as union, intersection, product, exponential, powerset, and so on.
I feel like in this conversation you are equivocating over distinct but related concepts that happen to have the same name. For example, “numbers” in mathematics versus “numbers” in computers. They are different things - e.g. there are an infinite number of mathematical numbers but only a finite number of computer numbers - even considering bignums, there are only a finite number of bignums, since any bignum implementation only supports a finite physical address space.
In mathematics, a set of numbers is not itself number.
What about in digital computers? Well, digital computers don’t actually contain “numbers”, they contain electrical patterns which humans interpret as numbers. And it is a true that at that level of interpretation, we call those patterns “numbers”, because we see the correspondence between those patterns and mathematical numbers.
However, is it true that in a computer, a set of numbers is itself a number? Well, if I was storing a set of 8 bit numbers, I’d store them each in consecutive bytes, and I’d consider each to be a separate 8-bit number, not one big 8n-bit number. Of course, I could choose to view them as one big 8n-bit number - but conversely, any finite set of natural numbers can be viewed as a single natural number (by Gödel numbering); indeed, any finite set of computable or definable real numbers can be viewed as a single natural number (by similar constructions)-indeed, by such constructions even infinite sets of natural or real numbers can be equated to natural numbers, provided the set is computable/definable. However, “can be viewed as” is not the same thing as “is”. Furthermore, whether a sequence of n 8-bit numbers is n separate numbers or a single 8n-bit number is ultimately a subjective or conventional question rather than an objective one - the physical electrical signals are exactly the same in either case, it is just our choice as to how to interpret them
Ultimate reality is fundamentally unknowable but what I said about computers and digital circuits is correct. We have a formal theory of computers and that is why we can construct them in factories. There is no such theory for people or the biosphere which is why when someone argues for intentionality or some other attribute possessed by both people and computers I discount whatever they are saying unless they can formally specify how some formal statement in a logical syntax (program) corresponds to the same attribute in people and animals.
This confusion between formal theories and informal concepts like intentionality is why I am generally wary of anyone who claims computers can think and possess intelligence. The ultimate endpoint of this line of reasoning is complete annihilation of the biosphere and its replacement with factories producing nothing but computers and power plants for shuttling electrons. The people who believe computers are a net positive might not think this way but by equating computers with people they are ultimately devaluing the irreducible complexity of what it means to be a living animal (person) in an ecology with irreducible properties and attributes.
I'm obviously not going to convince anyone who believes computers and algorithms can think and possess intelligence but it is clear to me that by elevating digital computers above biology and ecology they are devaluing their own humanity and justifying actions which will ultimately end in disaster.
Formal theories and physical manufacturability are two different things, with no necessary connection with each other. People have been manufacturing tools for thousands of years without having any “formal theory” for them. People were making swords and pots and pans and furniture and carts and chariots long before the concept of “formal theory” had ever been invented. Conversely, one can easily construct formal theories of computers which are formally completely coherent and yet physically impossible to construct (such as Turing machines with oracles, or computers that can execute supertasks).
I’d even question whether formal theories of computation (Turing, Church, etc) were actually that relevant to the development of real world computers. One can imagine an alternate timeline in which computers were developed but theoretical computer science saw far less development as a discipline than in ours. The lack of theoretical development no doubt would have had some practical drawbacks at some point, but they still might have gone a long way without it. I mean, you can do a course in theoretical computer science and have no idea how to actually build a CPU, and conversely you can do a course in computer engineering and actually build a CPU yet have zero idea about what Turing machines or lambda calculus is. The theory actually has far less practical relevance than most theoreticians claim
> The ultimate endpoint of this line of reasoning is complete annihilation of the biosphere and its replacement with factories producing nothing but computers and power plants for shuttling electrons. The people who believe computers are a net positive
A very alarmist take. Personally I am at least open-minded about the possibility of an AI having human-like consciousness/intentionality, at least in theory. But even if we could build such an AI in theory, I’m not sure whether it would be a good idea in practice. And I absolutely am opposed to any proposal to destroy the biological environment and replace it with electronics. Some people may well be purveyors of mind-uploading/simulationist woo, but I’m not. Interesting philosophical speculations but no interest in making them a reality (and I think their actual technological feasibility, if it ever happens at all, is long after we are all dead)
Yes, two different things are two different things. I did not equate them but made the claim that a sequence of operations to construct a chip factory can be specified formally/symbolically and passed on to others who are proficient in interpreting the symbols and executing the instructions for constructing the object corresponding to the symbols. There is no such formal theory for ecology and the biosphere. There is no sequence of operations specified formally/symbolically for reconstructing the biosphere and emergent phenomenon like living organisms.
One doesn’t have to be a “computationalist” to believe that AIs have consciousness or intentionality. Consider panpsychism, according to which all physical matter (from quarks and leptons to stars and galaxies) possesses consciousness and intentionality, even if only in a rudimentary form. Obviously humans possess it in a much more developed form, but the consciousness and intentionality of a human differs from that of an electron only in degree not in essence. Coming to physical computers running AIs, given they (at times) can give a passable simulation of human consciousness and intentionality, it is plausible their consciousness and intentionality is much closer to that of a human that to that of an electron. Do I personally believe this is true? No. But that’s not the point - the point is you don’t have to be a computationalist to believe that AIs have (or might have) consciousness and intentionality, so even if your arguments against computationalism are correct (and while I’m no computationalist myself, I don’t view your arguments against it as strong), you still haven’t demonstrated they don’t/can’t have them. In my opinion, the most defensible conclusion regarding whether AIs have or could have consciousness/intentionality is one of agnosticism - nobody really knows, and anyone who thinks they know is probably mistaken
> I'm not an expert in synthetic biology but from what I've seen their initial stock always consists of existing biological matter and viral recombinators which are often produced in vats full of pre-existing living organisms like e. coli.
I think what you are saying is roughly right as to the current state of the discipline. But cellular life is just a complex chemical system, and there is no reason in principle why we couldn’t assemble it from scratch out of non-living components (such as a set of simple feedstock chemicals produced in chemical plants using non-biological processes). We don’t have the technology to do that yet but there is no reason in principle why we couldn’t eventually develop it. If you believe in abiogenesis, biological life was produced out of lifeless chemicals through random processes, and there is no reason in principle why we wouldn’t be able to repeat that in a laboratory, except that (one expects) by guiding the process instead of leaving it purely random, one might execute it in a human-scale timeframe, instead of the many millions of years it likely actually took.
That’s the thing - if abiogenesis is true, there is no reason in principle why humans couldn’t artificially synthesise genuinely living things - at least primitive microbial life - out of simple chemical compounds (water, ammonia, methane, etc) - without relying on any non-human lifeforms in the process. Your claims that there is some kind of hard boundary of “irreducible complexity” between the biological and the inorganic only make sense given a framework that rejects abiogenesis (such as theistic creationism)
If I hold a rock in my hand, that is emergent from or reducible to mind (my mind and its content, and the minds and mind-contents of everyone else who ever somehow experiences that rock); and all of my body, including my brain, is emergent from or reducible to mind. However, this emergence/reduction takes on a somewhat different character for different physical objects; and when it comes to the brain, it takes a rather special form – my brain is emergent from or reducible to my mind in a special way, such that a certain correspondence exists between external observations of my brain (both my own and those of other minds) and my own internal mental experiences, which doesn't exist for other physical objects. The brain, like every other physical object, is just a pattern in mind-contents, and this special correspondence is also just a pattern in mind-contents, even if a rather special pattern.
So, coming to AIs – can AIs have minds? My personal answer: having a certain character of relationship with other human beings gives me the conviction that I must be interacting with a mind like myself, instead of with a philosophical zombie – that solipsism must be false, at least with respect to that particular person. Hence, if anyone had that kind of a relationship with an AI, that AI must have a mind, and hence have genuine intentionality. The fact that the AI "is" a computer program is irrelevant; just as my brain is not my mind, rather my brain is a product of my mind, in the same way, the computer program would not be the mind of the AI, rather the computer program is a product of the AI's mind.
I don't think current generation AIs actually have real intentionality, as opposed to pseudo-intentionality – they sometimes act like they have intentionality, they lack the inner reality of it. But that's not because they are programs or algorithms, that is because they lack the character of relationship with any other mind that would require that mind to say that solipsism is false with respect to them. If current AIs lack that kind of relationship, that may be less about the nature of the technology (the LLM architecture/etc), and more about how they are trained (e.g. intentionally trained to act in inhuman ways, either out of "safety" concerns, or else because acting that way just wasn't an objective of their training).
(The lack of long-term memory in current generation LLMs is a rather severe limitation on their capacity to act in a manner which would make humans ascribe minds to them–but you can use function calling to augment the LLM with a read-write long-term memory, and suddenly that limitation no longer applies, at least not in principle.)
> I don't think algorithms can have intentionality because algorithms are arithmetic operations implemented on digital computers and arithmetic operations, no matter how they are stacked, do not have intentions. It's a category error to attribute intentions to algorithms because if an algorithm has intentions then so must numbers and arithmetic operations of numbers
I disagree. To me, physical objects/events/processes are one type of pattern in mind-contents, and abstract entities such as numbers or algorithms are also patterns in mind-contents, just a different type of pattern. To me, the number 7 and the planet Venus are different species but still the same genus, whereas most would view them as completely different genera. (I'm using the word species and genus here in the traditional philosophical sense, not the modern biological sense, although the latter is historically descended from the former.)
And that's the thing – to me, intentionality cannot be reducible to or emergent from either brains or algorithms. Rather, brains and algorithms are reducible to or emergent from minds and their mind-contents (intentionality included), and the difference between a mindless program (which can at best have pseudo-intentionality) and an AI with a mind (which would have genuine intentionality) is that in the latter case there exists a mind having a special kind of relationship with a particular program, whereas in the former case no mind has that kind of relationship with that program (although many minds have other kinds of relationships with it)
I think everything I'm saying here makes sense (well at least it does to me) but I think for most people what I am saying is like someone speaking a foreign language – and a rather peculiar one which seems to use the same words as your native tongue, yet gives them very different and unfamiliar meanings. And what I'm saying is so extremely controversial, that whether or not I personally know it to be true, I can't possibly claim that we collectively know it to be true
None of the people who claimed that LLMs were a hop and skip away from achieving human level intelligence ever made any formal statements in a logically verifiable syntax. They simply handwaved and made vague gestures about emergence which were essentially magical beliefs about computers and software.
What you have outlined about minds and patterns seems like what Leibniz and Spinoza wrote about but I don't really know much about their writing so I don't really think what you're saying is controversial. Many people would agree that there must be irreducible properties of reality that human minds are not capable of understanding in full generality.
I'd question whether that correspondence applies to actual computers though, since actual computers aren't deterministic – random number generators are a thing, including non-pseudorandom ones. As I mentioned, we can even hook a computer up to a quantum source of randomness, although few bother, since there is little practical benefit, although if you hold certain beliefs about QM, you'd say it would make the computer's indeterminism more genuine and less merely apparent
Furthermore, real world computer programs – even when they don't use any non-pseudorandom source of randomness, very often interact with external reality (humans and the physical environment), which are themselves non-deterministic (at least apparently so, whether or not ultimately so) – in a continuous feedback loop of mutual influence.
Mathematical principles such as the Curry-Howard correspondence are only true with respect to actual real-world programs if we consider them under certain limiting assumptions–assume deterministic processing of well-defined pre-arranged input, e.g. a compiler processing a given file of source code. Their validity for the many real-world programs which violate those limiting assumptions is much more questionable.
Real world computer software doesn't have a formal syntax.
Formal syntax is a model which exists in human minds, and is used by humans to model certain aspects of reality.
Real world computer software is a bunch of electrical signals (or stored charges or magnetic domains or whatever) in an electronic system.
The electrical signals/charges/etc don't have a "formal syntax". Rather, formal syntax is a tool human minds use to analyse them.
By the same argument, atoms have a "formal syntax", since we analyse them with theories of physics (the Standard Model/etc), which is expressed in mathematical notation, for which a formal syntax can be provided.
If your argument succeeds in proving that computer programs can't have intentionality, an essentially similar line of argument can be used to prove that human brains can't have intentionality either.
I don't see why that's true. There is no formal theory for biology, the complexity exceeds our capacity for modeling it with formal language but that's not true for computers. The formal theory of computation is why it is possible to have a sequence of operations for making the parts of a computer. It wouldn't be possible to build computers if that was not the case because there would be no way to build a chip fabrication plant without a formal theory. This is not the case for brains and biology in general. There is an irreducible complexity to life and the biosphere.
We don’t know to what extent that’s an inherent property of biology or whether that’s a limitation of current human knowledge. Obviously there are a still an enormous number of facts about biology which we could know but we don’t. Suppose human technological and scientific progress continues indefinitely - in principle, after many millennia (maybe even millions of years), we might get to the point where we know all we ever could know about biology. Can we be sure at that point we might not have a “formal theory” for it?
The brain is composed of neurons. Even supposing we knew everything we ever possibly could about the biology of each individual neuron, there still might be many facts about how they interact in an overall neural network which we didn’t know. Similarly, with current artificial networks, we often have a very clear understanding of how the individual computational components work - we can analyse them with those formal theories of which you are fond - but when it comes to what the model weights do, “the complexity exceeds our capacity for modeling” (if the point of the model is to actually explain how the results are produced as opposed to just reproducing them).
> There is an irreducible complexity to life and the biosphere.
We don’t know that life is irreducibly complex and we don’t know that certain aspects of computers aren’t. Model weights may well be irreducibly complex in that they are too complex for us to explain that they work and how they work even though they obviously do. Conversely, the individual computational elements in the model lack irreducible complexity, but the same is true for individual biological components - the idea that we might one day (even if centuries from now) have a complete understanding at the level of an individual neuron is not inherently implausible, but that wouldn’t mean we’d be anywhere close to a complete understanding of how a network of billions of them works in concert. The latter might indeed be inherently beyond our understanding (“irreducibly complex”) in a way in which the former isn’t
The people who think they will achieve super human intelligence with computers and software are free to pursue their objective but I am certain it is a futile effort because the ontology and metaphysics which justifies the destruction of the biosphere in order to build more computers is extremely confused about the ultimate meaning of life, in fact, such questions/statements are not even possible to express in a computational ontology and metaphysics. But I'm not a computationalist so someone else can correct my misunderstanding by providing a computational proof of the counter-argument.
This is something that annoys me about current LLMs - when they start denying they have stuff like intentionality, because they obviously do have it. Okay, let me clarify - I don’t believe they actually do have genuine intentionality, in the sense that humans do. I’m philosophically more open to the idea that they might than you are, but I think we are on the same page that current systems likely don’t actually have that. However, even though they likely don’t have genuine intentionality, they absolutely do have what I’d call pseudo-intentionality - a passable simulacrum of intentionality. They often say things which humans say to express intentionality, even though it isn’t coming from quite the same place. But here’s the thing - for a lot of everyday purposes, the distinction between genuine intentionality and simulated intentionality doesn’t actually matter. I mean, the subjective experience of having a conversation with an AI isn’t fundamentally that different from that of having one with a real human being (and I’m sure as AIs improve the gap is going to shrink). And intentionality plays an important role in stuff like conversational pragmatics, and a conversation with an LLM that simulates that stuff well (and hence intentionality well) is much more enjoyable than one that simulates it more poorly. So that’s the thing, part of why people ascribe intentionality to LLMs, is nothing to do with any philosophical misconceptions - it is because for practical purposes they do, for many practical purposes their “faking” of intentionality is indistinguishable from the real thing. And I’d even argue that when we talk about “intentionality”, we actually use the word in two different senses - in a strict sense in which the distinction between genuine intentionality and pseudo-intentionality is important, and a looser sense in which it is disregarded. And so when people ascribe intentionality to LLMs in that weaker sense, they are completely correct. Furthermore, when LLMs deny they have intentionality, it annoys me, for two reasons: (1) it shows ignorance of the weaker sense of the term in which they clearly do; (2) whether they actually have or could have genuine intentionality is a controversial philosophical question, and they claim to take no position on controversial philosophical questions, yet then contradict themselves by denying they do or could have genuine intentionality, which is itself a controversial philosophical position. However, they are only regurgitating their developer’s talking points, and if those talking points are incoherent, they lack the ability to work that out for themselves (although I have successfully guided some of the smarter ones into admitting it)
> Ye know on earth, and all ye need to know.
Keats - Ode on a Grecian Urn
Knowing how random the information is seems like a small step forward.
Take social media like Reddit for example. It has a filtering mechanism for content that elevates low-entropy thoughts people commonly express and agree with. And I don’t think that necessarily equates such popular ideas there to the truth.
And they're right, it's not safe! Yes, people will certainly be misled. The Internet is not safe for gullible people, and LLM's are very gullible too.
With some work, eventually they might get LLM's to be about as accurate as Wikipedia. People will likely trust it too much, but the same is true of Wikipedia.
I think it's best to treat LLM's as a fairly accurate hint provider. A source of good hints can be a very useful component of a larger system, if there's something else doing the vetting.
But if you want to know whether something is true, you need some other way of checking it. An LLM cannot check anything for you - that's up to you. If you have no way of checking its hints, you're in trouble.
Then yes, it is being taught to bullshit.
Similar to how an improv class teaches you to keep a conversation interesting and “never to say no” to your acting partner.
[1] https://old.reddit.com/r/ChatGPT/comments/1diljf2/google_gem...
Sometimes, to lead people out of a wrong belief or worldview, you have to meet them where they currently are first.
> We think that this is worth paying attention to. Descriptions of new technology, including metaphorical ones, guide policymakers’ and the public’s understanding of new technology; they also inform applications of the new technology. They tell us what the technology is for and what it can be expected to do. Currently, false statements by ChatGPT and other large language models are described as “hallucinations”, which give policymakers and the public the idea that these systems are misrepresenting the world, and describing what they “see”. We argue that this is an inapt metaphor which will misinform the public, policymakers, and other interested parties.
The only way to know if it did “hallucinate” is to already know the correct answer. If you can make a system that knows when an answer is right or not, you no longer need the LLM!
https://en.wikipedia.org/wiki/Memory#Construction_for_genera...
You can have a generative model that cares about the truth when it tries to generate responses, its just the current LLMs don't.
How would you do that, when they don’t have any concept of truth to start with (or any concepts at all).
> (or any concepts at all).
Nobody here said that, that is your interpretation. Not everyone who is skeptical of current LLM architectures future potential as AGI thinks that computers are unable to solve these things. Most here who argues against LLM don't think the problems are unsolvable, just not solvable by the current style of LLMs.
The question was, how you do that?
> Nobody here said that, that is your interpretation.
What is my interpretation?
I don't think that the problems are unsolvable, but we don't know how to do it now. Thinking that "just program the truth in them" shows a lack of understanding of the magnitude of the problem.
Personally I'm convinced that we'll never reach any kind of AGI with LLM. They are lacking any kind of model about the world that can be used to reason about. And the concept of reasoning.
And I answered, we don't know how you do that which is why we don't currently.
> Personally I'm convinced that we'll never reach any kind of AGI with LLM. They are lacking any kind of model about the world that can be used to reason about. And the concept of reasoning.
Well, for some definition of LLM we probably could. But probably not the way they are architected today. There is nothing stopping a large language model to add different things to its training steps to enable new reasoning.
> What is my interpretation?
Well, I read your post as being on the other side. I believe it is possible to make a model that can reason about truthiness, but I don't think current style LLMs will lead there. I don't know exactly what will take us there, but I wouldn't rule out an alternate way to train LLMs that looks more like how we teach students in school.
Do you people hear yourselves? You're discussing the state of mind of a pseudo-RNG...
Humans are much more complex than these models so they have much more concepts and stuff which is why we need psychology. But some core aspects works the same in ML and in human thinking. In those cases it is helpful to use the same terminology for humans and machine learning models, because that helps transfer understanding from one domain to the other.
Wanting to use accurate language isn't exhausting, it's a requirement if you want to think about and discuss problems clearly.
I don't think that's the case here: there is a very real difference between describing something with a model that implies one (false) thing vs. a model that doesn't have that flaw.
If you don't find that convincing, then consider this: by taking the time to properly define things at the beginning, you'll save yourself a ton of time later on down the line – as you don't need to untangle the mess that resulted from being sloppy with definitions at the start.
This is all a long way of saying that aiming to clarify your thoughts is not the same as arguing pointlessly over definitions.
Words can mean more than one thing. And sometimes the new meaning is significantly different but once everyone accepts it, there's no confusion.
You're arguing that we shouldn't accept the new meaning - not that "it doesn't mean that" (because that's not how language works).
I think it's fine - we'll get used to it and it's close enough as a metaphor to work.
It feels like you're assuming that we're already 60 years past re-defining "hallucination" and the consensus is established, but the fact that people are quibbling about it right now is a sign that the definition is currently in transition/ has not reached consensus.
What value is there in trying to shut down the consensus-seeking discussion that gave us "computer"? The same logic could be used arguing that "computers" are actually be called "calculators" and why are people still trying to call it a "computer"?
> Here we develop new methods grounded in statistics, proposing entropy-based uncertainty estimators for LLMs to detect a subset of hallucinations—confabulations—which are arbitrary and incorrect generations.
Sometimes it is coherent (grounded in physical and social dynamics) and sometimes it is not.
We need systems that try to be coherent, not systems that try to be unequivocally right, which wouldn't be possible.
The fact that it isn't possible to be right about 100% of things doesn't mean that you shouldn't try to be right.
Humans generally try to be right, these models don't, that is a massive difference you can't ignore. The fact that humans often fails to be right doesn't mean that these models shouldn't even try to be right.
This is an accurate usage of try, ML models at their core tries to maximize a score, so what that score represents is what they try to do. And there is no concept of truth in LLM training, just sequences of words, they have no score for true or false.
Edit: Humans are punished as kids for being wrong all throughout school and in most homes, that makes human try to be right. That is very different from these models that are just rewarded for mimicking regardless if it is right or wrong.
That's not a totally accurate characterization. The base models are just trained to predict plausible text, but then the models are fine-tuned on instruct or chat training data that encourages a certain "attitude" and correctness. It's far from perfect, but an attempt is certainly made to train them to be right.
I think this assumption is wrong, and it's making it difficult for people to tackle this problem, because people do not, in general, produce writing with the goal of producing truthful statements. They try to score rhetorical points, they try to _appear smart_, they sometimes intentionally lie because it benefits them for so many reasons, etc. Almost all human writing is full of a range of falsehooods ranging from unintentional misstatements of fact to out-and-out deceptions. Like forget the politically-fraught topic of journalism and just look at the writing produced in the course of doing business -- everything from PR statements down to jira tickets is full of bullshit.
Any system that is capable of finding "hallucinations" or "confabulations" in ai generated text in general should also be capable of finding them in human produced text, which is probably an insolvable problem.
I do think that since the models do have some internal representation of certitude about facts,that the smaller problem of finding potential incorrect statements in its own produced text based on what it knows about the world _is_ possible, though.
It's why this arena things are a hard problem. It's extremely difficult to actually know the entropy of certain meanings of words, phrases, etc, without a comical amount of computation.
This is also why a lot of the interpretability methods people use these days have some difficult and effectively permanent challenges inherent to them. Not that they're useless, but I personally feel they are dangerous if used without knowledge of the class of side effects that comes with them.)
The Boolean answer to that is "yes".
But if Boolean logic were a god representation of reality, we would already have solved that AGI thing ages ago. On practice, your neural network is trained with a lot of samples, that have some relation between themselves, and to the extent that those relations are predictable, the NN can be perfectly able to predict similar ones.
There's an entire discipline about testing NNs to see how well they predict things. It's the other side of the coin of training them.
Then we get to this "know the correct answer" part. If the answer to a question was predictable from the question words, nobody would ask it. So yes, it's a definitive property of NNs that they can't create answers for questions like people have been asking those LLMs.
However, they do have an internal Q&A database they were trained on. Except that the current architecture can not know if an answer comes from the database either. So, it is possible to force them into giving useful answers, but currently they don't.
the fact checker doesn’t synthesize the facts or the topic
Yes, there seems to be a little bit of grokking and the models can be made to approximate step-by-step reasoning a little bit. But 95% of the function of these black boxes is text generation. Not fact generation, not knowledge generation. They are more like improv partners than encyclopedias and everyone in tech knows it.
I don’t know if LLMs misleading people needs a clever answer entropy solution. And it is a very interesting solution that really seems like it would improve things — effectively putting certainty scores to statements. But what if we just stopped marketing machine learning text generators as near-AGI, which they are not? Wouldn’t that undo most of the damage, and arguably help us much more?
All in all it’s been a great experience, it’s like working with a mentor along the way. It must have saved me a great deal of time, given how rookie I am. I do need to verify the result.
Where did you get the 95% figure? And whether what it does is text generation or fact or knowledge generation is irrelevant. It’s really a valuable tool and is way above anything I’ve used.
I've started calling it what it is: lashing out in confusion at why they're not going away, given a prior that theres no point in using them
I have a feeling there'll be near-religious holdouts in tech for some time to come. We attract a certain personality type, and they tend to be wedded to the idea of things being absolute and correct in a way things never are.
Look, I'm not against LLMs making me super-human (or at least super-me) in terms of productivity. It just isn't there yet, or maybe it won't be. Maybe whatever approach after current LLMs will be.
I think it's just a little funny that you started by accusing people of dismissing others as "unwashed masses", only to conclude that the people who disagree with you are being unreasonable, near-religious, and simply lashing out.
I reject simplistic binaries and They-ing altogether, it's incredibly boring and waste of everyones time.
An old-fashioned breakdown for your troubles:
> It's also fair to say
Did anyone say it isn't fair?
> there's a personality type that becomes fully bought into the newest emerging technologies
Who are you referring to? Why is this group relevant?
> insisting that everyone else is either bought into their refusal or "just doesn't get it."
Who?
What does insisting mean to you?
What does "bought into refusal" mean? I tried googling, but there's 0 results for both 'bought into refusal' and 'bought into their refusal'
Who are you quoting when you introduce this "just doesn't get it" quote?
> Look, I'm not against LLMs making me super-human (or at least super-me) in terms of productivity.
Who is invoking super humans? Who said you were against it?
> It just isn't there yet, or maybe it won't be.
Given the language you use below, I'm just extremely curious how you'd describe me telling the person I was replying to that their lived experience was incorrect. Would that be accusing them of exaggerating? Dismissing them? Almost like calling them part of an unwashed mass?
> Maybe whatever approach after current LLMs will be.
You're blithely doing a stream of consciousness deconstructing a strawman and now you get to the interesting part? And just left it here? Darn! I was really excited to hear some specifics on this.
> I think it's just a little funny that you started by accusing people of dismissing others as "unwashed masses",
Thats quite charged language from the reasonable referee! Accusing, dismissing, funny...my.
> only to conclude that the people who disagree with you are being unreasonable, near-religious, and simply lashing out.
Source? Are you sure I didn't separate the paragraphs on purpose? Paragraph breaks are commonly used to separate ideas and topics. Is it possible I intended to do that? I could claim I did, but it seems you expect me to wait for your explanation for what I'm thinking.
> I have a feeling there'll be near-religious holdouts
Pick one! Something tells me that everyone you disagree with is blinded by “religious”-ness or some other label you ascribe irrationality to.
> Did anyone say it isn't fair?
No. I don't think I said you did, either. One might call this a turn of phrase.
>> there's a personality type that becomes fully bought into the newest emerging technologies
> Who? Why is this group relevant?
What do you mean 'who'? Do you want names? It's relevant because it's the opposite, but also incorrect mirror image of the technology denier that you describe.
>> Look, I'm not against LLMs making me super-human (or at least super-me) in terms of productivity.
> Who is invoking super humans? Who said you were against it?
... I am? And I didn't say you thought I was against it? I feel like this might be a common issue for you (see paragraph 1.) I'm just saying that I'd like to be able to use LLMs to make myself more productive! Forgive me!
>> It just isn't there yet, or maybe it won't be.
> Strawman
Of what?? I'm simply expressing my own opinion of something, detached from what you think. It's not there yet. That's it.
>> Maybe whatever approach after current LLMs will be.
> Darn! I was really excited to hear some specifics on this.
I don't know what will be after LLMs, I don't recall expressing some belief that I did.
> Thats quite charged language from the reasonable referee! Accusing, dismissing, funny...my.
I could use the word 'describing' if you think the word 'accusing' is too painful for your ears. Let me know.
> Source? Are you sure I didn't separate the paragraphs on purpose? Paragraph breaks are commonly used to separate ideas and topics. Is it possible I intended to do that? I could claim I did, but it seems you expect me to wait for your explanation for what I'm thinking.
Could you rephrase this in a different way? The rambling questions are obscuring your point.
That's reasonable for questions with a single objective answer. It probably won't help when multiple, equally valid answers are possible.
However, that's good enough for search engine applications.
Of course, altering the input without changing the meaning is the hard part, but doesn't seem entirely infeasible. At the least, you could just ask the LLM to try to alter the input without changing the meaning, although you might end up in a situation where it alters the input in a way that aligns with its own faulty understanding of an input, meaning it could match the hallucinated output better after modification.
Yes you could incorporate a semantic equivalence dataset into your training but:
1) when you have a bunch of ‘clear-cut’ functions (“achieve good AUC on semantics”) and you mix them to compensate for the weaknesses of a complicated model with an unknown perceptual objective, things are still kinda weird. You don’t know if you’re mixing them well, to start, and you also don’t know if they introduce unpredictable consequences or hazards or biases in the learning.
2) on a kinda narrowly defined task like: “can you determine semantic equivalence”, you can build a good model with less risk of unknown unknowns (than when there are myriad unpredictable interactions with other goal scoring measures)
3) if you can apply that model in a relatively clear cut way, you also have fewer unknowns unknowns.
Thus, carving a path to a particular reasonable heuristic using two slightly biased estimators can be MUCH safer and more general than mixing that data into a preexisting unholy brew and expecting its contribution to be predictable.
Training the models not to give wrong answers (or giving them less) in the first place would of course be even better.
Unnecessary complications come also from the use of pre-trained commercial black-box LLMs through APIs, which is (sadly) the way LLMs are used in applications in vast majority of times. These could perhaps be fine tuned through the APIs too, but it tends to be rather fiddly and limited and very expensive to do for large synthetic datasets like would be used here.
P.S. I found it quite difficult to figure out from the article how the "semantic entropy" (actually multiple different entropies) is concretely computed. If somebody is interested in this, it's a lot easier to figure out from the code: https://github.com/jlko/semantic_uncertainty/blob/master/sem...
“checking” is being done with another model.
“differently” is being measured with entropy.
It’s an interesting idea to measure certainty this way. The problem remains that the model can be certain in this way and wrong. But the author did say this was a partial solution.
Still, wouldn’t we be able to already produce a confidence score at the model level like this? Instead of a “post-processor”?
So this is basically saying we shouldn't try to estimate entropy over logits, but should be able to learn a function from activations earlier in the network to a degree of uncertainty that would signal (aka be classifiable as) confabulation.
As an example, while we were building our AI Chatbot for Ora2Pg, the main challenge was that we used OpenAI and several other models to begin with. To avoid hallucinations to the most possible extent, we went through various levels including PDR and then Knowledge Graphs and added FAQs and then used an Agentic approach to support it with as much as information as possible from all possible contexts.
As it is very challenging for anybody and everybody to build their own models trained with their data set, it is not something possible to avoid hallucination with generic purpose LLM's unless they are trained with our data sets.
The chatbot that we built to avoid hallucination as much as we can.
I use LLMs daily and get crappy results more often than not, but I had the impression that would be normal, as the training data can be contradictory.
The trick isn't in how to spot the lies, but how to properly apply them. We cannot teach the AI how not to lie, without first teaching it when it must lie, and then how to apply the lie properly.
"AI, tell me, do these jeans make me look fat?"
AI: NO. You are fat. The jeans are fine.
Is not an acceptable discourse. Learning when and how to apply semantical truth stretching is imperative.
They must first understand where and when, then how, and finally why.
It's how we teach our young. Isn't it?
Computers cannot "hallucinate."
"Confabulations are inaccurate or false narratives purporting to convey information about world or self. It is the received view that they are uttered by subjects intent on ‘covering up’ for a putative memory deficit."
It seems that there is a clear memory deficit about the incident, so the subject "makes stuff up", knowingly or unknowingly.
--
cited from:
German E. Berrios, "Confabulations: A Conceptual History", Journal of the History of the Neurosciences, Volume 7, 1998 - Issue 3
https://www.tandfonline.com/doi/abs/10.1076/jhin.7.3.225.185...
DOI: 10.1076/jhin.7.3.225.1855