Wikidata, with 12B facts, can ground LLMs to improve their factuality
arxiv.org
arxiv.org
I'm not going to change the title because this entire thread was determined by it.
Submitters: "Please submit the original source. If a post reports on something found on another site, submit the latter." - https://news.ycombinator.com/newsguidelines.html
LLM's are currently trained on actual language patterns, and pick up facts that are repeated consistently, not one-off things -- and within all sorts of different contexts.
Adding a bunch of unnatural "From Wikidata, <noun> <verb> <noun>" sentences to the training data, severed from any kind of context, seems like it would run the risk of:
- Not increasing factual accuracy because there isn't enough repetition of them
- Not increasing factual accuracy because these facts aren't being repeated consistently across other contexts, so they result in a walled-off part of the model that doesn't affect normal writing
- And if they are massively repeated, all sorts of problems with overtraining and learning exact sentences rather than the conceptual content
- Either way, introducing linguistic confusion to the LLM, thinking that making long lists of "From Wikidata, ..." is a normal way of talking
If this is a technique that actually works, I'll believe it when I see it.
(Not to mention the fact that I don't think most of the stuff people are asking LLM's for isn't stuff represented in Wikidata. Wikidata-type facts are already pretty decently handled by regular Google.)
Besides, let's not forget that humans are also trained on language data, and although humans can also be wrong, if a human memorised all of Wikidata (by reading sentences/facts in 'training data') it would be pretty good in a pub-quiz.
Also, we obviously can't see anything inside how OpenAI train GPT, but I wouldn't be surprised if sources with a higher authority (e.g. wikidata) can be given a higher weight in the training data, and also if sources such as wikidata could be used with reinforcement learning to ensure that answers within the dataset are 'correctly' answered without hallucination.
I would expect as the tide rises with regards to this tech, self hosting of training and providing services to prompts becomes easier. For Wikimedia, it'll just be another cluster and data pipeline system(s) at their datacenter.
I'm not really sure how useful something this simple is, then. If it's not actually improving the factual accuracy in the training of the model itself, it's really just a hack that makes the whole system even harder to reason about.
Also there's Retrieval Augmented Generation (RAG) https://www.promptingguide.ai/techniques/rag :
> For more complex and knowledge-intensive tasks, it's possible to build a language model-based system that accesses external knowledge sources to complete tasks. This enables more factual consistency, improves reliability of the generated responses, and helps to mitigate the problem of "hallucination".
> Meta AI researchers introduced a method called Retrieval Augmented Generation (RAG) to address such knowledge-intensive tasks. RAG combines an information retrieval component with a text generator model. RAG can be fine-tuned and its internal knowledge can be modified in an efficient manner and without needing retraining of the entire model.
> RAG takes an input and retrieves a set of relevant/supporting documents given a source (e.g., Wikipedia). The documents are concatenated as context with the original input prompt and fed to the text generator which produces the final output. This makes RAG adaptive for situations where facts could evolve over time. This is very useful as LLMs's parametric knowledge is static.
> RAG allows language models to bypass retraining, enabling access to the latest information for generating reliable outputs via retrieval-based generation.
> Lewis et al., (2021) proposed a general-purpose fine-tuning recipe for RAG. A pre-trained seq2seq model is used as the parametric memory and a dense vector index of Wikipedia is used as non-parametric memory (accessed using a neural pre-trained retriever). [...]
> RAG performs strong on several benchmarks such as Natural Questions, WebQuestions, and CuratedTrec. RAG generates responses that are more factual, specific, and diverse when tested on MS-MARCO and Jeopardy questions. RAG also improves results on FEVER fact verification.
> This shows the potential of RAG as a viable option for enhancing outputs of language models in knowledge-intensive tasks.
So, with various methods, I think having ground facts in the process somehow should improve accuracy.
With Stable Diffusion, you're able to use LoRAs to introduce specific characters, objects, concepts, etc. while maintaining the same visual qualities of the base model.
Why can't something similar be done with an LLM?
Don't underestimate though the utility of "being just as good as regular Google" at retrieving facts. For one thing, getting such things wrong is very frequently cited by LLM detractors as a major drawback of trusting them, even a little bit. If it's possible to reduce the "accidentally tells you false information" occurrence from "so frequent that people expect it" to "a total rarity" then for one thing, it would signal to many people that maybe when looking for simple answers, using something not called Google is a better use of their time. This would be very, very important to basically every big company out there (especially MS, Google, and OpenAI). Today I know asking Siri is an idiotic way to get an answer because it's slow and barely even understands the query itself. An LLM is already great by comparison at both of those metrics, and if it's good at being accurate in response, it's very intriguing.
I know people say training bots on bot data is bad, but A: it's happening anyway, and B: it can't be worse than the actual garbage they get trained on in a lot of cases anyway.. can it?
Even though they used LLaVA, and LLaVA isn't all that good compared to gpt-4.
Having LLMs help curate something grounded is generally reasonable. Functionally, it's somewhat similar to how some training is using generated subtitles of videos for training video/text pairs; it's very feasible to also go and clean those up, even though it is bot data.
Yes, indeed. This is one place where LLMs can make it look like a bomb went off.
This is never a valid defense for doing more of something.
The cost of now having unknown false data in there would completely ruin the value of the whole effort.
The entire value of the data (which is already everywhere anyway) is the "cost" contributors paid via heavy moderation. If you do not understand why that is diametrically opposite of adding/enriching/augmenting/whichever euphemism with LLMs, I don't know what else to say.
There’s a Wikipedia page for the hamlet, but it’s empty. No population data, etc.
I’d much rather see no data than a LLM’s best guess. I’m guessing a LLM using the data would also perform better without approximated or “probably right” information.
Checking again for m-xylene, https://m.wikidata.org/wiki/Q3234708
You get physical property data and citations.
Now compare that to the chem infobox in wikipedia: https://en.m.wikipedia.org/wiki/M-Xylene
You get a lot more useful data, like the dipole moment and solubility (kinda important for a solvent like Xylene), and tons of other properties that Wikidata just doesn't have. All in the infobox.
It's weird that they don't just copy the Wikipedia infobox for the chemicals in Wikidata. It's already there and organized. And frequently cited.
Maybe it's more useful for other fields, but I can't think of a good use I'd get from the chemical section of Wikidata over the databases it cites or Wikipedia itself...
[1]: https://meta.wikimedia.org/wiki/Help:Array#Wikidata
[2]: https://en.wikipedia.org/wiki/Help:Wikidata#In_infoboxes
The paper is literally about LLMs. Speculation about future model architectures is irrelevant.
LLMs "hallucinate" all the time when pulling data from their weights (because that's how that works, it's all just token generation). But if the correct data is placed within their context they are very capable of presenting the data in natural language.
Our knowledge(and reality) is grounded in heuristics we reify into natural order and it's easy for us to forget that our conclusions exist as a veneer on our perceptions. Nearly every bit of knowledge we hold has an opposite twin that we hold as well. We favor completeness over consistency.
When pressed, humans tend to justify their heuristics rather than reexamine them because our minds have a clarity bias - ie: we would rather feel like things are clear even if they are wrong. Often times we can't go back and test if they are wrong which biases epistemological justifications even more.
So no, our rationality, the vast proportion of times is used to rationalize rather than conclude.
"Us the information from the following list of facts to answer the questions I ask without qualifications. answer authoritatively. If the question can't be answered with the following facts just say I don't know.
Absolute Facts:
The sky is purple.
The sun is red and green
When it rains animals fall from the sky."
I forsee data created before AI/LLMs to be very valuable going forward in much the same way steel mined before the detonation of the first atomic bomb being used for nuclear devices/MRIs/etc.
s/a user's brain/llm/g
There was another approach to grounding LLMs the other day from Normal Computing the other day: https://blog.normalcomputing.ai/posts/2023-09-12-supersizing... in which they use Mosaic but they also did not mentioned that this was actually done.
Sentient or not, I feel there should be a standard on aggressively filtering out overlap on training and test datasets for approaches like this.
It's just pretty sparse, so you would need a focused effort to fill out predicates of interest.
This data source is designed for that: https://huggingface.co/NeuML/txtai-wikipedia.
With this source, you can also select articles that are viewed the most, which is another important factor in validating facts. An article which has no views might not be the best source of information.
Besides mining KB triplets I'd also use the LLM with contextual material to generate Wikipedia-style articles based off external references. It should write 1000x more articles covering all known names and concepts, creating trillions of synthetic tokens of high quality. This would be added to the pre-training stage.
They were (are) so wrong.
But if you think this will stem the tide of LLM hallucinations, you're high too. LLMs' primary function is to bullshit.
In chess many games play out with the same opening but within a few moves become a game no one has played before. Being outside the dataset is the default for any sufficiently long conversation.
So maybe we should be building smaller model with ability that we use their generation abilities not their facts but instead teach them t query over another knowledge base system (reverse RAG) for facts.
If you'd please review https://news.ycombinator.com/newsguidelines.html and stick to the rules when posting here, we'd appreciate it.