Training LLMs to generate text with citations via fine-grained rewards
arxiv.org
arxiv.org
"When possible cite your sources. Use the following custom html format <citation>{document_id}†{node_id}</citation>. Very important!"
An LLM is based on the idea that there is some very complicated function
y(t) = f(y(t-1), y(t-2), ..., y(t-K))
for any sequence of text tokens t(1), y(2), ... and some huge context size K.We fit the model by attempting to find an approximation of f() with a low badness score (called "loss"). We also don't usually want the absolute lowest badness score, because there is always some tradeoff between between minimizing loss on the specific test data that we happen to have, and preserving the ability to generalize to content outside of the test data.
The technique here is an improvement to that process of finding a good approximation of f(). The specific technique here is to adjust the loss to mean something more specific than "error rate on predicting the next word in a sequence."
The entire basis of the principle is that the loss function controls what the model learns. If we want it to produce more sentences that look a certain way, we introduce a reward for generating sentences that look that way. If we want to avoid certain kinds of outputs, we introduce a penalty for generating those kinds of outputs. From there, it's just gradient descent. Better outputs produce lower losses, so the model starts producing better outputs, without us having to really know or care about what the model is doing internally.
The technique of RLHF is similar along those lines. We discourage the model to hallucinate by having a human review its output and report when the model hallucinates, so we can impose a penalty for hallucination, and thereby shift the model output in the direction of not-hallucinating after many such rounds of "learning".
Is it a fundamental shift in how LLMs work? No. Could it possibly lead to improvements? Yes.
Maybe! I'm not an AI researcher or mathematician, so I don't know if anyone has pursued this idea. The problem might be that any such structure is intractably complicated to describe within the limits of human understanding.
> How do you know that there are no discontinuous regions or that you are not stuck in a local minima?
Are you talking about f() getting stuck in some kind of bad region while generating text, or about the optimization process itself?
The answer is the same in both cases: we don't.
Regarding text generation, we've seen plenty of examples where specific prompts result in pathological output. Although I'm not sure if newer models have that problem.
Regarding optimization, there is absolutely no guarantee that we have found a global minimum. Some interesting research has been done on the "loss landscape" of these giant neural network models, and my understanding is that they are messy and complicated. Keep in mind that the training data is part of the loss function! Finding a local minimum on the training data might just result in overfitting.
Do you have any recommended reading for this? It sounds like a super interesting area of research.
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, Tom Goldstein. 2018. "Visualizing the Loss Landscape of Neural Nets".
;)
> ... And now, since you are the father of writing, your affection for it has made you describe its effects as the opposite of what they really are. In fact, it will introduce forgetfulness into the soul of those who learn it: they will not practice using their memory because they will put their trust in writing, which is external and depends on signs that belong to others, instead of trying to remember from the inside, completely on their own. You have not discovered a potion for remembering, but for reminding; you provide your students with the appearance of wisdom, not with its reality. Your invention will enable them to hear many things without being properly taught, and they will imagine that they have come to know much while for the most part they will know nothing. And they will be difficult to get along with, since they will merely appear to be wise instead of really being so.”
Note that the criticism is mostly spot on!
I spend some time in an internet forum where memory athletes hang out and you'd be surprised at what the human ability to remember can actually be.
This is from Phaedrus by Plato.
There are a few extra observations in the following text that nicely seems applicable to LLMs
https://conversational-leadership.net/myth-of-thamus-and-the...
For those interested in an alternate method that doesn't depend on a LLM, check out this article: https://neuml.hashnode.dev/build-rag-pipelines-with-txtai
Disclaimer: I'm the primary author of txtai.
I will say though that hopefully you'll consider applying that policy equally to all. Because many VC-backed and large companies basically post press releases and they trend without issue.
I'm a single person open-source project. But it's your site, I'll respect your request and not post moving forward.
I appreciate all you do in keeping the site up and running along with the dedication to ensuring it has high-quality content.