Teach your LLM to answer with facts, not fiction
blog.myscale.com
blog.myscale.com
"What does Charmander evolve into?"
"What does the spell 'avada kedavra' do?"
"What is the Sindarin word for 'friend'?"
"What are the names of Santa's reindeer?"
"Where did Robin Hood live?"
"Where did Achilles die?"
These are all 'factual questions' you can find answers to from reputable sources like Wikipedia. Google displays 'fact boxes' for several of them. Wolfram Alpha provides answers for three of them. The answers to some of these questions are part of what passes for 'general knowledge' in some societies.It's no surprise that LLMs trained on human writings produce text which claims things that aren't true are facts. Humans do that all the time.
There are well attested reputable sources that will tell you Abraham Lincoln was a vampire hunter, others that say he was a Lego Master Builder, and others still will tell you that among his notable quotes is "Party on dudes - be excellent to each other". So what's an LLM to do when it's trying to extend a paragraph of information about Abraham Lincoln?
When an LLM is suggesting what might come next in a piece of text... it doesn't know if it's supposed to guess a probable word from a Wikipedia article, an Onion article, a Project Gutenberg manuscript, or an Archive Of Our Own fanfic. So you get a bit of all that.
Because of elision.
"[Homer wrote] that Achilles died of an arrow in the heel"
This is why the Wiener Kreis taught to use protocolar statements: "<There> and <at that time> <that individual> witnessed <that fact>".
Edit: oh, by the way, in case of interest: https://plato.stanford.edu/entries/vienna-circle/
> "protocoli(s|z)ed"
the use of '-ize' is (a graecism) indicated by the OED as International English, as opposed to British, American etc. In fact, some call International English "British spelling with -ize" - it is not exactly that but close. One exception is 'analyse', but that is because linguists compromised on the "difficult" original 'analysize'.)
It could be. I cannot bring to mind the rules for doubling right now. They both occur, 'protocolar' much more often. I will correct my original post.
The obvious start seems to be having separate fiction and nonfiction LLMs and not training the nonfiction ones on Archive Of Our Own. People also end up confused about the truth when nobody points out the difference between fiction and nonfiction.
The obvious answer then is to tell it to make sure that what it's finally outputting is really part of the "real" API, but I think it's safe to say there's some technical hitch there, as it's safe to say OpenAI probably spent quite a lot of energy trying to solve the code hallucinations, and ultimately was unable to do so. I'd guess that the more you restrict its recombination ability, the more you end up with it inappropriately (and incorrectly) just regurgitating large chunks of its training input verbatim. Basically it becomes more like a keyword hunting search engine, and less like a generative LLM.
How about an economics textbook, or an article in the economist? "A history of the english speaking peoples" by Winston Churchill?
If we restrict to "ground truth we feel very sure about" it feels like available training data might be quite small.
And anyway - answers to all my ‘fictional facts’ questions above can be sourced from Wikipedia - there’s tons of made up stuff on there.
Hint: how many stadiums are filled with people standing up to recite a newspaper article about a battle in the War of 1812?
This is true of base LLM models that are just trained on missing-word prediction on the training corpus, but one of the main points of RLHF[1] is to tune this model to make these kind of inferences the way a human would expect. For example if you asked an untuned model to write a poem in the style of ... etc., a valid internet response might be "hmm no thanks, you go first", you need to steer the model away from replying like this.
I'm not saying it's perfect, but it's wrong to say e.g. GPT-4 has had no information about the difference between a good and bad response and is just generating internet-like text at random, the big players have made progress on this already.
[1] https://en.wikipedia.org/wiki/Reinforcement_learning_from_hu...
Reinforcement learning trains them that question and answer sessions contain answers which statistically correlate with factual statements in their broader learning corpus.
When formulating answers, this leads them to formulate answers that reflect the factual information on which they were trained.
My point is that the source data contains a far muddier range of information than just unarguable facts.
We largely want LLM based Q&A bots to answer questions about fictional or mythical characters in their own terms. As I said, those questions above all have reasonably ‘correct’ answers.
The fact that from all that LLMs do as well as they do is remarkable. But it also seems like it requires us to assume that LLMs are capable of a remarkable degree of cultural nuance, media literacy and contextual awareness for them to figure out the different authorship, salience, trustworthiness, agenda, biases, and assumptions of all the gigareams of text they’ve ingested.
LLMs are very good at inferring context, so that only really applies if you’re using an un-RLHFed base model with no context given
So no LLM knows what it's supposed to do. If you prefer, you could say it only ever has one goal: to generate a sequence of tokens which are jointly the most probable to occur along with the prompt tokens, given such probabilities in a historical corpus.
This imitates knowledge, goal-directness, "inferring context" etc. without doing any of those things. Consider what the aim of knowing, goal-directness, inferring , etc. is --- it is never "consistency with a historical text corpus".
For knowing: that beliefs correspond to the way the world is; for goal-directness that one's acts+desires can realise changes; for 'inferring context': that one is sensitive to reasons to speak outside of what is literally spoken.
LLMs are never sensitive to reasons to speak outside of what has been spoken.
RLHF is the difference between GPT-3.5 and ChatGPT, and it's the whole reason why LLMs are suddenly such a big deal. ChatGPT demonstrated that it's possible to give language models a goal beyond just "complete most likely next word" and that they can actually be somewhat competent at achieving those goals despite not being explicitly trained for them.
Well (1) it doesn't achieve goals, since a "goal" is observer-relative. We have goals, the LLM has a formal optimisation objective which gives it the appearence of goal-directed behaviour (in a similar way, eg., that it appears pens want to fall when dropped).
And (2), reading your "goal" here even in observer-relative ways, I don't think there's much evidence of this. These models are "trained" on everything ever written, include all of the internet and basically all digitised book. I don't see any evidence of much generalisation -- if you can find it by google, then the LLM has it stored compressed (ie., the "weights").
The innovation in LLMs is being able to compute `max P(answer|prompt, historical_corpus)` for increasingly longer prompts --- there's no innovation in goal-directed behaviour.
That's VC propangada to disguise the fact that LLMs are mostly an innovation in copyright laundering.
(2) "I don't see any evidence of much generalisation" Seriously? So when I tell ChatGPT to rewrite a paragraph in the style of Shakespeare and it does it, despite never being trained to do that, never seeing the source or target paragraph before, and having no information other than my text prompt and its past training, that's not evidence of generalization? And that's only one of millions of different possible tasks that the same model excels at, despite being trained on nothing but a bunch of unstructured text and a few examples indicating its goal should be to follow instructions given in the prompt text. Up until a couple years ago this level of flexibility in a machine learning model would have been considered science fiction by nearly everyone, and now it's "[not] evidence of much generalization". Okay.
Is the child a genius or are they just reading out of a textbook? Can the toddler really compose a sonata or did they just press play on the piano keyboard?
(2) This is indeed the power of interpolating between the data points of "everything ever written in human history" as digitised and compressed by ChatGPT.
If you have 1 billion circles of radii 0 to 1, it isn't generalisation for the machine to produce one with a radii 0.0000100003000001, ie., one not in the set but a mere interpolation of points within it.
It would be expensive, but imagining "reversing" ChatGPT from it's output to the sources which made a non-trivial difference to generating that output.
So the function there is: response -> verbatim text in the training corpus.
Then, maybe, "bolded" by how much each paragraph would "make a difference" to its output.
What you'd find is thousands of pages: all Shakespeare ever written, all papers about Shakespeare, all books about Shakespeare; and so on.
Then when it applied the bolding, and summarising it a little, the trick would be revealed: it would be apparent how a naive statistical interpolation between sequences of characters could produce the effect.
ChatGPT exists because of ebooks and social media: without it, it could do almost nothing. That is, the appearance of these capacities is strictly derivative of the work of a billion people who had them.
Without vast, unimaginable, amounts of work produced on Shakespeare this system wouldnt work. It's just a copyright laundering system. All the school essays on reddit, all the forum posts; all of usenet. All pdfs, all digitised works. All academic papers.
Is this generalisation? Is this a system which starts with little and makes a lot?
Or is it a system which is more like a child reading from a textbook? Ie., making a haphazard ability to repeat what's already written.
The size of the weights of a modern LLM are sufficient to compress everything ever written in human history: and that's exactly what they do.
If there's truly a difference between "a capacity [and] an apparent capacity" then you should be able to point out what that difference actually is in practice. A child pressing play on a piano can only play one song. A LLM composing poems can compose billions upon billions of unique, never-before-seen poems about every conceivable topic. Whether under the hood it does that by "interpolating numbers in n-dimensional spaces" or "some incomprehensible arrangement of neurons linked together" or some other, yet to be invented process doesn't matter if the result is the same. The fact that you can explain how something works doesn't make it less real.
It does always amaze me that we trained LLMs on a dump of the internet and then people are shocked that they're about as trustworthy as a random web page.
There also exists the consideration of allographemical contextualization, the nature of relevance, pragmatics, conjunct identification of context, semantics. To be honest the linguistics side alone is vast. Knowledge and cognition however. . . A whole other ballgame. But the only tool we have to really get down to the bottom of how knowledge works is language, it's to epistemological pursuit what math is to physics.
While GPT is super impressive and can do a lot of quasi-brute-force things, we're only finding now the rudiments of the machined intelligence paradigm, and it will behoove any reader to brush up on their classics, true pursuants of philosophy and many order logic are about to be in high demand if I had to reckon.
And when answering a question, unambiguously specify this fictional context, or at least indicate that it might be fiction if unsure.
"You need to believe in things that aren't true. How else can they become?"
- "Hogfather" by Terry Pratchettcharmander -> pokemon -> fiction avada kedavra -> harry potter -> fiction sindarin -> ??? -> infer( fiction or nonfiction) Robin Hood -> disambiguation -> ask(user input-> do you mean?) ...
This just seems like a categorization and data annotation problem, which I would assume a bunch of projects are trying to solve like this one.
wait why is this implied to not be black and white? Charmeleon is the only correct answer.
In general facts are not the answer to Hallucinations. You can't possibly have every fact for every situation. The true solution to Hallucinations is figuring out how to make a model say 'I don't know"
(Issue is, now some are convinced that people in general would do the same and just blurt out the feedforward output of their "internal neural network", as opposed to having built knowledge in a loop of critical evaluation.)
And why would that be? "Hallucination" means "erratic wandering", implying one is lost - similarly to "delirium" (maetaphor using the plough) and "error". Part of the idea is that of "instead of witnessing the correct, reporting the false" - a very ancient, traditional idea, and akin to the concept of "intelligence" (intus-legere).
"Confabulation" means locutor and interlocutor are talking, exchanging narrations.
but I am not sure - provisionally - that it can be a good idea to relate strictly human neurology to ANNs, if based on phenomena as opposed to structural issues. You do not have that problem when staying with natural language.
Still have to figure a measurement unit.
Second, it isn’t even necessarily better to have fewer lies if those few lies are more subtle. Plenty of propaganda works by twisting facts and using misleading statements. Perhaps the worst offenders won’t even have any outright falsehoods at all.
There's a categorical difference between knowing a fact, and looking up a fact. When you know a fact you can recognize it in a situation where you wouldn't know to look it up, and you'd know to utilize it in a larger solution rather than simply parrot it when specifically asked about it.
Databases of facts have and will still have their place, but that is absolutely not the solution to LLM telling apart fact from truth. They have to innately have this in their model. I don't believe the nature of LLM is to hallucinate. It's instead a side effect of how we train them. We train them to guess, to be close, but not to be correct necessarily. And why is it a surprise that's precisely what they do?
Also LLM are too small in order to be accurate. They're tiny. GPT4 is roughly 40 times smaller than a human brain. And GPT4 is very large compared to GPT-3, and GPT-3 is very large compared to LLaMA 2.
We'll need for hardware to catch up so we can scale things up pragmatically and see what happens to their ability to grasp facts. But also architectural changes, of course.
Thoughout this comment you speak about LLMs as-if they're animals, or real physical objects. An LLM is a formal model which is just to generate a sequence of tokens maximally probabilistically consistent with a corpus of historical text.
A digital machine running a LLM program is a physical object which necessarily generates text based on "guessing" because that's the algorithm it's running. LLMs are "guessing algorithms", all of Machine Learning is -- it is dumb brute-force analysis of conditional probability.
> GPT4 is roughly 40 times smaller than a human brain
This doesn't make any sense. GPT4 is an abstract algorithm with no "size". The brain has 10^{big number} cells, and GPT4 can be specified with a single real number. Is that the comparison to make? No, both comparisons are incoherent.
A physical device running GPT4 can be given a "size", but it would again have nothing to do with a brain.
LLMs arent living things where we can "measure their size" and "train them to know, rather than to guess". They are just the equation, `max P(answer|propmt, historical_corpus)`
A machine running GPT4 is just an electrical device generating text according to the rule given above. There is no sense of "training it to do something other than guesswork", and no sense of "size"
What they are literally doing is guessing the next word, a word a time but doing it really really well and making statistically average output over a very large number of inputs.
There is no distinction between understanding "the" vs "a" and telling me 1+1=3. It is all token generation.
For what that concerns us here: LLMs will never learn to fact-check anything. They'll blindly regurgitate the facts they have been "taught", but never consider or evaluate "the paper cited for this fact on wikipedia is a bunch of bullshit".
Any attempt to use them to produce "facts" is ultimately just folly, in the same way Google's attempt to do so with it's search engine index is.
Nor do people, though! This is setting the bar way too high.
The whole point to having edited reference sources like "encyclopedias" is that so that we can rely on the expertise of the editors in lieu of having to develop the expertise ourselves[1].
No, an LLM that simply knows a priori (via prompt hacking) which sources are trustworthy would be absolutely comparable to the way an educated-but-non-expert human approaches sources.
[1] Which is a chicken and egg problem anyway. Everyone starts with edited reference sources as tutorial material. Quite frankly everyone starts learning with wikipedia.
No. If these things are claimed to be sources of truth, then the bar needs to be that high.
It is precisely because people don't fact-check that the bar has to be so high.
That's a strawman, though. No service, nor human, "claims to be a source of truth" in the kind of profound sense you seem to be using. It stops, everywhere, at "Wikipedia (or whatever) said it and I trust it".
The only way to get access to deeper expertise is to (1) BE an expert and (2) engage in an discussion with another.
GPT doesn't do math correctly but it also doesn't just memorize it.
What is most probable is not always what is most correct or most accurate.
Why is that necessary? Why have an LLM guess where the facts are?
Put all of that data in a place where it's normalized and ready to vector search.
This has not been my experience. Did you create any benchmarks as a part of this project?
My startup has a product for lawyers that uses RAG to answer legal queries (https://lawlight.ai/). We have a disclaimer that "... (we) do not guarantee the accuracy of answers. You are responsible for reviewing the cited case law and drawing your own independent conclusions."
(This works within the specific context—lawyers are domain experts; and they are supposed to read through all cases they cite in court anyway.)
* I dislike the term "hallucinations." By definition LLMs hallucinate. It's just that much (or most) of the time, the hallucinations reflect reality.
I thought it involved prompting the LLM to write SQL code to query a knowledge base of documents, and index into them, so that you'd know where to look in the original documents for your authoritative answer. So it would be a meta-search agent.
But apparently, they intend the queried documents to feed back into training the LLM? That's just gasoline on a dumpster fire.
The LLM layer seems completely unnecessary. Why do you have a schema that requires an LLM to decide which column to query (which is the LLM's only unique value in this proposal)? Why are you not normalizing into a single column?
Oh, we have something similar: perplexity.ai
It provides a number of sources after prompting its textual result.
Transformer-based LLMs are interesting because they are such good version of autocomplete that they can, for example, complete a news article about scientists discovering unicorns, using just the first sentence (this was one of the first public demonstrations of GPT-2). But fundamentally they are still just auto-complete.
Assistant: ...
User: ...
Assistant: ...
And the output is stopped when they start generating the equivalent of "User: " and the reins are handed back to you.
This isn't a problem, autocomplete at the level of "what would a person say next" is outrageously powerful, but it is how they're working afaik.
I.e. synthetic professionals. (Reliable things. Problem solvers.)
Problem: LLMs are not search engines. They extrapolate, interpolate, and approximate (so-called “hallucinations”) so they can always produce somewhat-plausible text completions.
Solution: Create a search engine so good at returning relevant results that even an LLM can make use of it… then go to significant lengths to plug that search engine into the LLM, preventing people from reading the search results directly.
Why not simply give people access to the search engine‽ People know how to use search engines!! This is the fifth time I've seen an article like this, and I'm still… baffled. It's https://xkcd.com/2021/ all over again.
They enable voice-based interfaces to be practical for normal users for the first time since you really can talk to them in a convincing way.
Translating user input into a set of well-defined commands seems like a better use than searching data to me.
Half the power of LLMs as they currently exist is that they can often extract the intention of the user's question in a way that search engines usually can't, allowing them to provide a more useful answer or at least point the user in the right direction.
Perhaps it would make sense for search engines to utilize LLMs to perform this query extraction and suggest more appropriate search terms, engaging conversational interaction only if the suggestions are wrong and the LLM requires further clarification.
I don’t know about that. Google is pretty good at including “similar” questions that others have asked to what I have queried and often that’s exactly what I needed.
The whole debate is revolving around hot air, because nobody knows whether the other person is talking about the same thing as themselves.
If you define "facts" as "things actually stored in the LLM's weights", then research shows it is possible to determine if an output is a "fact" or not.
Although looking on arxiv I found a paper saying it doesn't work (https://arxiv.org/pdf/2307.00175.pdf) so maybe not.
(Oh, what a matter: * all the epistemological debate - hardly a deterministic solution; * the fact that we cannot train a function approximator through supervised learning; * the challenge of unsupervised learning; * the scientific and teleological problem that, if we have an ANN find a solution, what we may want is to go "Ok black box, now teach us how you do it to expand our knowledge (not just our dumb capabilities)...")
You cannot solve the problem of discrimination (in non trivial cases of true and false, of good and bad) through a deterministic solution, as the epistemological debate did not solve the general problem. You cannot train a function approximator (e.g. an ANN) as a Discriminator through supervised learning, because the problem remains such for human judgement. Creating a Discriminator through unsupervised learning, I'd like to see how one would frame a proposal; and anyway, if we could create reliable filters, the main question - as usual for a progressive approach to AI - would be to have the oracle in the system teach us instead that knowledge that what could not achieve with good old thinking.
For the first puzzle she had to put the cities Tokyo, Paris, LA, and Brisbane (where we live) in order. Using colours, since she can’t read yet!
Then we had to discuss why those cities were in that order (answer below).
For the video, I figured I would doctor a Google search to land on a page explaining the answer. The actual article I found that mostly worked was about 12 results down.
Then I tried Bing’s Chat instead. It spat out the correct answer in 2 sentences. Deus ex machina indeed! (The four cities are the Summer Olympic hosts, 2020-2032).
So I disagree that it’s impossible. And I can absolutely see the value - asking “why are these cities in this order?” is a real question, like “what might be causing the squeaking sound when my car brakes?” or “what was the movie Audrey Hepburn made with the photographer?”
Search Engines aren’t great for those kind of questions - just google “What’s a good brownie recipe?”. LLMs can give the user exactly what they want, or even prompt for the additional extra context.
Not that they will always be correct; hence the flaw.
From the looks of it, they're just advertising a SQL extension that adds vectorization and vector search. Further, it looks like the only thing that the LLM is doing here is deciding which column to run the vector search on. Why is that even necessary? Why are you not pre-processing "vector'd" columns into a normalized format to query against?
They're basically adding an unnecessary LLM step to what amounts to a vector search. In fact, the LLM is essentially blindly deciding which column is the best column to pull an answer from.
-----
EDIT: Just struck me how terrifyingly dangerous this blog is. Really tired of seeing this crap in the LLM community.
The basic premise of this blog is "give an LLM complete access to your database. Let it decide how and where it should pull data from". This is basically useless without talking about how you prevent the LLM from pulling data from places you don't want it to.
A far better and safer approach remains to push your relevant fields to a separate place for your LLM. In the spirit of this blog, you should just index to a new table. More realistically, you should just put this in a vector store.
Using GPT4 and Code Interpreter, I have asked it to write a function and test it, given some inputs and expected outputs. The function returned different values when it tested it, but it lied and said it worked as expected.
You need to read the code and the test outputs yourself. Or maybe have it write an automated test?
Despite this, it seems quite promising. I expect that in a year or two, some IDE's will come with a useful pair programming feature.
The "test" was a print statement, not a unit test. There wasn't a failure message. It had to read the output and compare it to the expected value I gave it.
It claimed it got a different result. I guess it didn't really read the result because it strongly expected something else?
If you use an assertEquals() that loudly complains, maybe it's less likely to do this? I haven't seen it ignore stack traces.
This article sounds like an idea I had independent not too long ago, but with a different goal:
LLMs are great at natural language comprehension, but also have a lot of neurons dedicated to factoids. Using neurons that way is really inefficient, can we split the "language" capability from the "knowledge" capability and have the former just look things up in a database?
My question was more about reducing the size of the network rather than reducing hallucinations, but it's still a separately updatable knowledge resource.
(The answer may actually be "no"; I don't study this professionally, but technical jargon is kinda both factual domain knowledge and also linguistic comprehension, which is why Oracle isn't competing with Starbucks for Java beans).
> LLMs are great at natural language comprehension, but also have a lot of neurons dedicated to factoids. Using neurons that way is really inefficient, can we split the "language" capability from the "knowledge" capability and have the former just look things up in a database?
I think the beauty of LLMs is exactly that all we need to do is to feed them raw text --- the hope, I guess, had been that the models will be able to develop human-like insights by learning to understand and "speak" languages on its own. Introducing "feature engineering" (e.g. the distinction as you suggested) would defeat that goal.
They do, that's how tokens are selected. Locally run models or non-chat ones from openai can return the probabilities and you can do things like modify or filter them.
The superpower has been the ability to synthesize output from very disparate training sources, and any answer for "I don't know" would come from needing to synthesize disparate training sources.
- The usual knowledge cutoff warning
- An explanation of what a hallucination is
If (in a separate conversation) I give GPT-3.5 the exact prompt that explains what an LLM is, I get gaslighted instead. GPT-3.5 attempts to tell me that LLM stands for "Legal Master of Laws". Then it gives the knowledge cutoff warning, and then the same correct explanation Myscale got.
The rest of this article appears to be trying to turn GPT into a frontend for search engines. I don't know why people keep trying to do this.
https://law.pepperdine.edu/blog/posts/llm-versus-jd-degree.h...
> In other words, a hallucination is an error in (or a false) perception of something real or concrete.
an llm has no "perception" it doesn't "believe" or "think" that the answers it provides are "correct" or "true" or even "false". It's just autocompleting strings with the most probably next words.
If we keep treating these things as if they're sentient entities that "want" to provide "correct" answers we're going to keep tripping over our false assumptions about their answers.
That being said, what you've described is different to how a human first learns and many never grow beyond in what way?
I guess this is a pretty unsolvable problem with current architectures. There's just no concrete "confidence" value. I mean, an LLM will give you a probable value for what confidence could be given the words preceding it, but that's an entirely different thing
I think a LLM interface to Wikipedia could be useful, at least I imagine it would.
It should be fairly trivial for it to tell you how often its straight up lied about something.
"Is the sky blue?"
* Generally, yes. If you ask a 5 year old, the answer is yes.
* Is the sky blue right now? Maybe, maybe not. You need to look outside. Even then, you might have wild fire haze. Is it still blue? Is it orange? When does blue become orange?
* Is the sky blue in Blade Runner? Doesn't really seem like it.
-----
Further, who are you talking to? Is this trivia night where your best guess is better than no guess? Is this a scientific panel? Do you have alternative options? Do those alternative options align with your opinion? If you're wrong, how wrong are you?
By what mechanism would it "lie" to you? How is it even capable of lying in this sense of the word:
https://en.wiktionary.org/wiki/lie#Verb_2
> To give false information intentionally with intent to deceive.
I guess it could "lie" in other senses, but calling that lying is not really adding clarity to the situation.
Luckily my use cases have a manual check built in but even a proxy for confidence would be amazing.
1) better filter the training data
2) design better retrieval and reranking algorithms
3) when context information is provided, make it use the sources and cite the sources (use extractive QA to highlight which part of the source is relevant. This is the type of hallucinations that we should focus on as we can compare the generated result and the context to detect the hallucinations)
4) make the llm break down its reasoning into small steps that can be validated inidividually (COT, PAL)
There are some research on how to manipulate the logits during decoding to make the generated text satisfy certain contraints. I suspect that we can use these techniques to make the LLM stick to the provided context. - Controllable Text Generation with Language Constraints
- Classifiers are Better Experts for Controllable Text Generation
- Stay on topic with Classifier-Free Guidance