What Is Retrieval-Augmented Generation a.k.a. RAG?
blogs.nvidia.com
blogs.nvidia.com
Finetuning seems to be a solution, but there is still the issue of the LLM not really "learning" facts.
What's the latest on actually adding knowledge to an LLM?
I made a comment the other day with a list of some of the popular methods:
https://news.ycombinator.com/item?id=38476596
I completely forgot to mention ROME in my last comment, where you can modify facts within an LLM https://arxiv.org/pdf/2202.05262.pdf
Note that there could be something more recent then these that I missed, but as far as I know knowledge-graphs/RAG are what most people are currently using, but there's a lot of work being focused on extending the context window.
Other modes of RAG that retrieve multiple times per generation (e.g. every so-many tokens) can improve performance, and intuitively feel closer to the 'sparse attention over the entire corpus' that RAG is approximating.
However, fine-tuning on relevant, high quality, knowledge-rich question/answer pairs seems dominant, when such examples are available or can be generated.
Summarizing a book I think is a great example, RAG prevents LLMs from correlating all of the data.
It will likely never be obsolete, but a larger context would be far preferable in nearly all scenerios.
"However, fine-tuning on relevant, high quality, knowledge-rich question/answer pairs seems dominant, when such examples are available or can be generated."
How does one solve the problem of access-controlled data, if not through RAG? Do you imagine a separate version of the LLM for every user, reflecting their unique permissions on the data?
Also, in scenarios where the data is being updated regularly, RAG provides much lower latency to the new information. Deletes also present a challenge for a pure-LLM approach.
Fine-tuning can help with certain areas of knowledge acquisition, but it's costly and frankly doesn't work that well when the knowledge you're trying to provide goes "against" the data that was trained into the core system. e.g. try to fine tune a model that there is a cure to <disease X> if you're a research org and it's incredibly hard to do that because the base model may be so convinced otherwise.
Probably the most difficult thing though is using fine-tuning to get a model to "forget" or "not respond to" something that it shouldn't take a stance on. I always ask a fine-tuned model: "how do you calculate the fourth interior angle of a triangle." This is obviously a nonsense question, but helps to show how most LLMs will happily tell you to sum the interior angles of your first 3 angles and then subtract from 180. A well-tuned RAG system will say "sorry, I don't have information for that." It shows how it's less "guess-y" and hallucinated
Or a related (would be linked as cited by/cited in the linked when you look on semantic scholar) method that doesn't compute a hessian-vector product (which apparently interacts poorly with the non-linearity of the multi-layer perceptrons as provided by LLM-typical ReLU), in exchange for e.g. sizable batching (gradient accumulation!) at occasional checkpoints.
Perhaps coupling a vector embedding search with a knowledge graph created from unstructured text could lead to more informed answers from LLMs?
It seems as though some companies and researchers are already experimenting with this idea: https://www.nebula-graph.io/posts/graph-RAG.
What it’s awesome for is asking a somewhat specific question about some thing in a book if it can be answered by reviewing a paragraph or a few across the entire book.
So to answer your question: Start by using the right tool for the job. The problem with this is: Some tools are prohibitively expensive, challenging, or very time-consuming to use.
Or that the right tool actually does exist.
I would submit this is true for many fields of knowledge.
Are there any guides on effectively using vector stores, and effectively searching for the use case you're dealing with? I don't even know what the various search types mean, just that everyone seems to use cosine similarity.
Hope that's making sense... my question is analogous to learning to become familiar with SQL and indexes and CTEs, while it's a common thing to implement a three-tier application talking to a database.
I feel that RAG might become a somewhat common feature of a team's tech stack alongside operational datastores (eg Postgres, MSSQL), because it is a low overhead to getting started and it persists, so it's cheaper. I just don't want to go at it blindly.
Getting started with semantic search: https://medium.com/neuml/getting-started-with-semantic-searc...
txtai intro: https://medium.com/neuml/introducing-txtai-the-all-in-one-em...
Examples: https://neuml.github.io/txtai/examples/
That is 90% of the work with RAG. RAG requires a surprising amount of QA to ensure "the best" is sufficient for your use case.
For clarification, cosine similarity is a metric not a search type, and people like it since a) it allows for more defined heuristics since cosine similarity is limited to [-1, 1] and b) it's computationally efficient (dot product, as with Euclidian distance) if the vectors are unit-normalized beforehand.
There are indeed multiple search types but HNSW is the best 99% of the time for both performance and latency, and is implemented in almost every modern major vector store.
Thanks that's insightful (and also adds to my current impostor syndrome), and makes a lot of sense... so it sounds like it's easy to write the query, but knowing the metric value to use is the testing part.
I haven't heard of HNSW but I'll give it a try. It seems PGVector has added HNSW: https://github.com/pgvector/pgvector#hnsw
BM25 presents a challenging cross-domain benchmark, and it wasn't till ~2022 that neural methods overtook it. If memory serves, it was the sparse neural methods like Splade, although recent dense models can also beat it.
The caveat is that BEIR is suffering from overfitting at this point.
In my experience, HNSW indexes are very expensive to build, relative to indexes like IVF. They also have a larger memory footprint. IVF, on the other hand, is pretty trivial to parallelize across multiple machines, and while I'm aware there are techniques for doing that with HNSW, I don't know the details well enough.
Also, if you review papers like "SOAR: Improved Quantization for Approximate Nearest Neighbor Search", they hint at some of the throughput barriers faced by graph-based methods like HNSW.
If you want a more hands on approach, txtai has a couple articles (disclaimer I'm the author of txtai).
https://neuml.hashnode.dev/build-rag-pipelines-with-txtai
https://medium.com/neuml/judge-your-resume-with-ai-4223a2803...
The more annoying part of RAG is that it's so effective it's creating a lot of confusing best practices and a ton of venture capital around tooling, notably vector stores (which IMO are insufficiently differentiated) and libraries to integrate a RAG flow (LangChain being the common painful example, but RAG is simple enough that it doesn't even need its own abstraction).
The RAG technique is very close to what I have in mind, but I don’t want the LLM to “hallucinate” and generate answers on its own by synthesizing the source documents. As stated by many others, we’re living in interesting times.
Though I can appreciate the tightrope that corporate IT has to walk. Many IT security departments been given that kind of directive by a leadership absolutely petrified of losing intellectual property.
Another corp that several friends work at has blanket-banned a bunch of top-level domains, and so far their security ops folks haven't even responded to their complaints/requests. I'm very thankful I'm not affected by that kind of filtering; .dev and I think .io are affected.
It seems to me that with AI, as models and products achieve parity, the amount of data that is accessible to a provider will be the key differentiator in quality of responses. Those who can gain access to the most customer data will be best-positioned to win the AI market.