Finetuning seems to be a solution, but there is still the issue of the LLM not really "learning" facts.
What's the latest on actually adding knowledge to an LLM?
Finetuning seems to be a solution, but there is still the issue of the LLM not really "learning" facts.
What's the latest on actually adding knowledge to an LLM?
Other modes of RAG that retrieve multiple times per generation (e.g. every so-many tokens) can improve performance, and intuitively feel closer to the 'sparse attention over the entire corpus' that RAG is approximating.
However, fine-tuning on relevant, high quality, knowledge-rich question/answer pairs seems dominant, when such examples are available or can be generated.
Summarizing a book I think is a great example, RAG prevents LLMs from correlating all of the data.
It will likely never be obsolete, but a larger context would be far preferable in nearly all scenerios.
"However, fine-tuning on relevant, high quality, knowledge-rich question/answer pairs seems dominant, when such examples are available or can be generated."
How does one solve the problem of access-controlled data, if not through RAG? Do you imagine a separate version of the LLM for every user, reflecting their unique permissions on the data?
Also, in scenarios where the data is being updated regularly, RAG provides much lower latency to the new information. Deletes also present a challenge for a pure-LLM approach.
I made a comment the other day with a list of some of the popular methods:
https://news.ycombinator.com/item?id=38476596
I completely forgot to mention ROME in my last comment, where you can modify facts within an LLM https://arxiv.org/pdf/2202.05262.pdf
Note that there could be something more recent then these that I missed, but as far as I know knowledge-graphs/RAG are what most people are currently using, but there's a lot of work being focused on extending the context window.
Fine-tuning can help with certain areas of knowledge acquisition, but it's costly and frankly doesn't work that well when the knowledge you're trying to provide goes "against" the data that was trained into the core system. e.g. try to fine tune a model that there is a cure to <disease X> if you're a research org and it's incredibly hard to do that because the base model may be so convinced otherwise.
Probably the most difficult thing though is using fine-tuning to get a model to "forget" or "not respond to" something that it shouldn't take a stance on. I always ask a fine-tuned model: "how do you calculate the fourth interior angle of a triangle." This is obviously a nonsense question, but helps to show how most LLMs will happily tell you to sum the interior angles of your first 3 angles and then subtract from 180. A well-tuned RAG system will say "sorry, I don't have information for that." It shows how it's less "guess-y" and hallucinated
Or a related (would be linked as cited by/cited in the linked when you look on semantic scholar) method that doesn't compute a hessian-vector product (which apparently interacts poorly with the non-linearity of the multi-layer perceptrons as provided by LLM-typical ReLU), in exchange for e.g. sizable batching (gradient accumulation!) at occasional checkpoints.
Perhaps coupling a vector embedding search with a knowledge graph created from unstructured text could lead to more informed answers from LLMs?
It seems as though some companies and researchers are already experimenting with this idea: https://www.nebula-graph.io/posts/graph-RAG.
What it’s awesome for is asking a somewhat specific question about some thing in a book if it can be answered by reviewing a paragraph or a few across the entire book.
So to answer your question: Start by using the right tool for the job. The problem with this is: Some tools are prohibitively expensive, challenging, or very time-consuming to use.
Or that the right tool actually does exist.
I would submit this is true for many fields of knowledge.