A Comprehensive Guide for Building Rag-Based LLM Applications
github.com
github.com
This post is mostly about plumbing. It's probably the right way to do it if it needs to be scaled. But for learning, it obscures what is essentially simple stuff going on behind the scenes.
[0] - https://github.com/emptycrown/llama-hub/blob/main/llama_hub/...
```
def queryPostgres(client, string)
return client.query(string)
```If you know what you want to build, building from scratch is easier than you think. If you're tinkering on the weekend, then maybe the frameworks are helpful.
Arguably, for many use cases (e.g. searching through a document with ~200 passages), loading embeddings in memory and running a simple linear search would be fast enough.
[1]- https://towardsdatascience.com/in-context-learning-approache...
Full disclosure, I just joined Zilliz this week as a Dev Advocate.
A vector storage could help in reduce the time it takes to retrieve the most similar hit. I used faiss as a local vector store quite a bit to retrieve vectors fast. Though I had 1.5 million vectors to work through.
An in memory index is about as good as it gets for a single node performance, and fitting that many vectors into memory on a single machine is easy.
You're likely to get better results from vector-based semantic search though, just because it takes you beyond needing exact matches on search terms.
I've found that internal enterprise projects tend to be very keyword based, and vector search often produces weird, head-scratcher results that users hate - whereas term-based search does a better job of capturing the right terms, if you do the proper synonym/abbreviation expansions.
That said, I use them both, usually with vector search as a fallback after the initial keyword-based RAG pass
We then just fetch up to the vectors related to a customer's schema in memory (largest is ~200MB) and run cosine similarity in a few ms in Go (handwritten, ~25 lines of code), and then we've got out top N things to place in our prompt.
Primitive? You betcha. Works extremely well for our entire customer base? Yup. You definitely don't need a Vector DB unless you have an enormous amount of vectors. For us it means having to run our own Redis clusters, but we know how to do that, and so we don't need to involve another vendor.
You definitely do need information retrieval. It just shouldn't be limited to vector dbs. Unfortunately vector db companies and the VCs that back them have flooded the internet with propaganda suggesting vector db is the only choice. https://colinharman.substack.com/p/beware-tunnel-vision-in-a...
For most serious use cases, you'll have far too much data to fit into 1 (or several) inference contexts.
We have been building RAG systems in production for a few months and have been tinkering with different strategies to get the most performance out of these pipelines. As others have pointed out, vector database may not be the right strategy for every problem. Similarly there are things like lost in the middle problems (https://arxiv.org/abs/2307.03172) that one may have to deal with. We put together our learnings building and optimizing these pipelines in a post at https://llmstack.ai/blog/retrieval-augmented-generation.
https://github.com/trypromptly/LLMStack is a low-code platform we open-sourced recently that ships these RAG pipelines out of the box with some app templates if anyone wants to try them out.
This will be the case when you're exposing an interface to end users that they can submit arbitrary queries to - such as "how do I turn off reverse breaking".
By converting the user's query to vectors before sending it to your vector store, you're getting at the user's actual intent behind their words - which can help you retrieve more accurate context to feed to your LLM when asking it to perform a chat completion, for example.
This is also important if you're dealing with proprietary or non-public data that a search engine can't see. Context-specific natural language queries are well suited to vector databases.
We wrote up a guide with examples here: https://www.pinecone.io/learn/retrieval-augmented-generation...
And we've got several example notebooks you can run end to end using our free-tier here: https://docs.pinecone.io/page/examples
Cosine similarity across vectors isn't enough here, but when combined with an LLM we get the right behavior. As you mention, without the vector store reducing the size of data we pass to the LLM, hallucinations happen more often. It's a balancing act.
The other nasty one to consider is when people write "how do I not turn off reverse breaking". Again, a comparison will show that as very similar to your input, but it's really the opposite. And so if implementers aren't careful to account for that, they've now got a nasty subtle bug on their hands.
The original paper proposing this technique can be found here: https://arxiv.org/pdf/2212.10496.pdf
Some of my concerns:
1) Is sentence embedding using an off-the-shelf embedding model going to capture the "meaning" of my logs? My answer is "probably not". For example, if a portion of my logs is in this format
timestamp_start,ClassName,FunctionName,timestamp_end
Will I be able to get meaningful embeddings that satisfy a query such as "what components in my system exhibited an anomalously high latency lately?" (this is just an example among many different queries I’d have)Based on the little I know, it seems to me off-the-shelf embeddings wouldn't be able to match the embedding of my query with the embeddings for the relevant log lines, given the complexity of this task.
2) Is it going to be even feasible (cost/performance-wise) to use embeddings when one has a firehose of data coming through, or is it better suited for a mostly-static corpus of data (e.g. your typical corporate documentation or product catalog)?
I know that I can achieve something similar with a Code Interpreter-like approach, so in theory I could build a multi-step reasoning agent that starting from my query and the data would try to (1) discover the schema and then (2) crunch the data to try to get to my answer, but I don't know how scalable this approach would effectively be.
From your questions it looks like you are only interesting in the R part. RAG implies the retrieval step is then used to augment a user prompt.
To answer 1, a good heuristic would be "can a human reasonably familiar with the terminology answer questions about the meaning?" If a human would need extra info to make sense of your data then so would an LLM.
This is where RAG typically comes in. For example if you had documentation about ClassName and FunctionName, a retrieval model might be able to find the most likely candidates based on a file containing full definitions of these classes and function, then pass that info into the LLM appended to your query.
For 2: It depends if the fire house is the query or the data. If you have queries coming in very quickly, then you might be able to if your firehose doesn't have too much volume since you can batch requests and get responses fairly quickly.
If the fire hose is the data going into the vector DB then you might have some difficultly inserting and indexing the data fast enough.
What RAG here is doing is using embeddings and a vector store to identify close pieces of information, for example "in this django project add a textfield" will be very close to documentation in the django docs that say "textfield", and it will then add that to the prompt so the LLM has the relevant docs in its context.
The problem is that you'll need a heuristic to identify at least "potentially anomalous" and even then you'll still have to make sure there's enough context for it to know "is this a normal daily fluctuation".
A multi-step agent is definitely what you want, you could have it build an SQL query itself, for example "was there any high latency requests yesterday?" it may identify it should filter the time, possibly design the query to determine what is "high".
---
At the moment I don't think it's well suited to identifying when the "latency is abnormally high". However if you have some other system/human identify heuristics to feed to the LLM, it may then be able to do at least answer the query.
I was trying to understand if there is an opportunity to introduce some of this technology to solve “anomaly detection” on large amount of structured data, where anomaly might be an incredibly overloaded term (it might imply a performance regression, a security issue, etc). That is a business need I have today.
It seems that what is possible today is an assistant that can aid a user to get to these answers faster (by, for instance, suggesting a SQL query based on the schema, etc). Again, roughly the equivalent of what Code Interpreter does, just without the local environment limitations.
- In the cold start section, a couple of the synthetic_data responses say 'context does not provide info..'
- It's strange that retrieval_score would decrease while quality_score increases at the higher chunk sizes. Could this just be that the retrieved chunk is starting to be larger than the reference?
- Gpt 3.5 pricing looks out of date, it's currently $0.0015 for input for the 4k model
- Interesting that pricing needs to be shown on a log scale. Gpt-4 is 46x more expensive than llama 2 70B for ~.3 score increase. Training a simple classifier seems like a great way to handle this.
- I wonder how stable the quality_score assessment is given the exact same configuration. I guess the score differences between falcon-180b, llama-2-70b and gpt-3.5 are insignificant?
Is there a similarly comprehensive deep dive into chunking methods anywhere? Especially for queries that require multiple chunks to answer at all - producing more relevant chunks would have a massive impact on response quality I imagine.
I do wonder, is there some bias in quality measures? Using GPT 4 to evaluate GPT 4's output? https://www.linkedin.com/feed/update/urn:li:activity:7103398...
https://www.anyscale.com/blog/a-comprehensive-guide-for-buil...