I want flexible queries, not RAG
win-vector.com
win-vector.com
I've worked with various search backends for about 20 years. People treat vector search like magic pixie dust but the reality is that it's not that great unless you heavily tune your models to your use cases. A well tuned manually crafted query goes a long way.
Pretty much any system I've built over the last few years, the best way to think about search is about building a search context that includes anything relevant to answering the user's question. The user's direct input is only a small part of that. In the case of mobile systems, the user entered query is actually typically a very minor part of it. People type two or three letters and then expect magic to happen. Vector search is completely useless in situations like that. Why does search on mobile work anyway? Because of everything else we know to create a query (user location, time zone, locale, past searches, preferences, etc.)
RAG isn't any different. It's just search where the search results are post processed by an LLM with whatever the user typed. The better the query and retrieval, the better the result. The LLM can't rescue a poorly tuned search. But it can dig through a massive result of search results and extract key points.
And there is value in that.
I generally agree with the OP: in a search context I often don't want that last LLM summarisation step. I want embeddings generated from my search query for document lookup, and I want to see the original sources. Phind did (does?) something like this with citations, which is at least an improvement.
The LLM step does provide a benefit, as it can synthesize a tight and concise answer to the user's question, instead of returning a list of source documents to sift through manually.
Oh how I want information. Oh how I don’t want your newlsetter modal to obscure that information.
But it's very telling that users are supportive of Google's agenda of not sending you to people's actual web sites because web sites have gotten so bad. Europeans will deny it but I think those cookie banners broke the dam so that now it is normalized that you can pop up four different modals asking for an email address.
Most of my search now happens in llm. Usually I want information, not a webpage. I do understand the writer's gripes with search, and it should have been fixed a long time ago.
I believe that what most people miss is that embeddings are the product of the model.
Embeddings have been getting so much better with the new models and so embedding-based search improved too.
Embeddings are coordinates in a world representation that "organizes" information so that location in that world matters and serve as a way to differentiate meaning.
In other worlds, word2vec was a simple and poor world representation. OpenAI embeddings are an astounding world representations coordinates.
The generation is a showy gimmick.
So why aren't we separating the useful bit out? My sneaking suspicion is that we can't. It's a package deal in that there are no two parts, it's just one big soup of free-associating some text with other text, the stochastic parrot.
LLM do not understand. They generate associated text that looks like an answer to the question. In order to separate the two parts, we'd need LLMs that understand.
That, apparently, is a lot harder.
If you wanted to create an llm that significantly improved on this you would probably need to make sure to massively clean up / reorganize your training data so that you always provide sufficient context and the llm is disincentived from baking in "facts". But I'm not sure this is tractable to do currently at the scale of data needed.
--
[0] - Well, most of it anyway; I bet the training set contains some amount of purely random text, for example technical articles discussing RNGs and showcasing their output. Some amount of noise is unavoidable.
(I’ve dubbed this sort of thing “nonstandard ML" since, like you, I have a fondness for thinking of unorthodox solutions that seem plausible.)
It really isn't? You can tell it to output in a JSON structure (or some other format) of your choice and it will, with high reliability. You control the output.
Honestly I wonder if the people who criticize LLM's have made a serious attempt to use them for anything
I mean, this is provably false. Have you tried to use LLMs to generate structured JSON output? Not only do all LLMs suck at reliably following a schema, you need to use all kinds of "forcing" to make sure the output is actually JSON anyway. By "forcing" I mean either (1) multi-shot prompting: "no, not like that," if the output isn't valid-ish JSON; or (2) literally stripping out—or rejecting—illegal tokens (which is what llama.cpp does[1][2]). And even with all of that, you still won't really have a production-ready pipeline in the general case.
[1] https://github.com/ggerganov/llama.cpp/issues/1300
[2] this is cutely called "constraining" a decoder; what it actually is is correcting a very clear stochastic deficiency in LLMs
Yeah it's worked about fifty thousand times for me without issues in the past few months for several NLP production pipelines.
Guidance (the industry term for "constraining" the model output) is only there to ensure the output follows a particular grammar. If you need JSON to fit a particular schema or format, then you can always validate it. In case of validation failure you can always pass the JSON and the validation result back to the LLM for it to correct it.
FWIW I'm using GPT4o not Llama, I've tried Llama for local tasks and found it pretty lacking in comparison to GPT.
Consider the possibility that at least some of the criticisms of LLMs are a result of serious attempts to use them.
I would not want, for example, a version of Copilot that slowly gets better at helping me with the task I'm working on and then suddenly and unpredictably reverts back to zero just because the k8s pod running the instance that I had been interacting with got recycled. Consistently mediocre behavior would be preferable to the AI equivalent of the pair programming equivalent of speed dating.
They may not be all that stable, either. I would not just assume that knowing how to interpret the attention heads in the base model of GPT-4 tells you anything about what the corresponding attention heads are doing in GPT-4t or GPT-4o.
* The prompt engineering side is a black art.
* Generation is pretty amazing but less useful than it first seems.
* The open source and proprietary tooling around managing and interrogating datasets for NLP stuff is way ahead of where it was last time I looked maybe a decade ago
* It looks to me that with appropriate thinking about indexing then there's a lot of potential here. For example if I have a question answer set, I can index the question and which answers point to it. Get the LLM to identify the type of question I've just asked, and then use my corresponding corpus of answers to provide something useful based on the appropriate parts of the answer corpus.
That last bit is not the automatic panacea that the flashy and somewhat gimicky emergent properties of the generative side supply, but it seems like it can get good traction on some quite difficult problems quickly, and actually quite well using not too many local computing resources.
The moment we say "an AI", we anthropomorphize the model. The narrative gets even more derailed when we say "Large Language Model". No LLM actually contains a grammar or defined words. An LLM can't think logically or objectively about subjects the way a person can, either.
I propose that we instead call them "Large Text Models". An LTM is a model made from training a neutral net with text. Sure, the text itself was written with language, but no part of the training process or its result does anything about it.
The really cool trick that an LTM does is to construct continuations whose content just happens to be indistinguishable from language. This is accomplished because the intentions of the original human writer (that were encoded into the original text dataset) did consistently follow language rules. The problem is that the original writer did not encode the truthiness or falsiness of future LTM continuations. Truth and lie are written the same, and that ambiguity lives in the LTM forever.
Search is often (but not exclusively) performed using cosine similarity over semantic vectors. Such vectors are produced using embedding models, which represent the meaning of the document via an arbitrary length vector called an embedding, and 768 is a common vector size for this.
You calculate the embedding for all documents in your database ahead of time (during insertion), and you calculate the embedding for the user's query, then search for documents closest to the query using a similarity metric of choice such as cosine.
Nothing prevents you from serving documents found this way directly, instead of using them to generate answers. Part of the Google search pipeline involves something like this. Many full-text search products also do this (Algolia is such an example https://www.algolia.com/blog/ai/what-is-vector-search/)
LLMs and generation are used to synthesize the answer and tie it back into the question using a layer of soft judgment based on the LLMs prior knowledge. This does work out great in some contexts, less so in others, as you pointed out. But these components aren't coupled in any way.
It was trained as a chat-bot so that's all the current training can do. If you want to use it, you must hack something useful out of the chat bot interface.
It was trained as a chat bot because that was the impressive thing that got them investors. Useful applications need a lot more context to awe people, and context takes time and work to create.
So, now that they've got money, are those companies that created those LLM chat-bots making a useful next generation engines behind the scenes? Well, absolutely not! The same situation applies on every researching round, and they need to show impressive results right now to keep having money.
(And now I wonder... Why do VC investors exist again?)
> they need to show impressive results right now to keep having money
Sure. You do as little as possible to make as much money as possible. This is a fundamental of commerce/human existence. But, at some point, it will end with everyone's models performing similarly, with free models catching up. The concept of sustained "impressive results" will eventually require actual "reasoning systems". They'll use all that accumulated wealth, from what you maybe perceive as low hanging fruit, to tackle it. I think it must be assumed that these AI companies are intentionally working toward that, especially since it's the stated goal of many of them. I think you must assume that these people are smart, and they can see the reality of their own systems.
This seems like a misunderstanding of what RAG is. RAG is not used to try to anchor to reality a general LLM by somehow making it come up with sources and links. RAG is a technology to augment search engines with vector search and, yes, a natural language interface. This concerns, typically, "small' search engines indexing a specific corpus. It lets them retrieve documents or document fragments that do not contain the terms in the search query, but that are conceptually similar (according to the encoder used).
RAG isn't a cure for ChatGPT's hallucinations, at all. It's a tool to improve and go past inverted indexes.
> RAG is not used to try to anchor to reality a general LLM by somehow making it come up with sources and links.
but that definitely is one particular use of RAG, i.e. to limit some potential hallucinations by grounding it in data provided in the prompt.
RAG is vector search first. It encodes the query, finds nearest vectors in the vector database, retrieves the fragments attached to those vectors, and then sends those vectors to the LLM for it to summarize them.
A general LLM like Gemini or Claude or ChatGPT first produces an answer to a question, based on its training. This doesn't involve searching any external source at that point. Then after that answer is produced, the LLM can try to find sources that match what it has come up with.
People, we have to stop talking about what we know as though it's all there is. Don't confuse our knowledge for understanding. Understanding only comes from repeadly trying to prove our understandings wrong and learning how things truly are.
Two common alternatives to vector search:
1. Ask the LLM to identify key terms in the user's question and use those terms with a regular full-text search engine
2. If the content you are answering questions about is short enough - an employee handbook for example - just jam the whole thing in the context. Claude 3 supports 200,000 tokens and Gemini Pro 1.5 supports a million so this can actually be pretty effective.
It’s bloody useful if you can’t cram your entire proprietary code base into a prompt.
ID: 1
URL: https://example.com/test.html
Text: lksdjflkdsjlksjkl
And then tell it to use the ID to link to the page.
> but that definitely is one particular use of RAG
One use of RAG doesn't imply that all uses of RAG are for grounding LLM hallucinations.
Additionally, I disdain when fallacies come up in conversation when it doesn't appear someone is making an argument in bad faith. We all use logical fallacies by accident from time to time. A good use of calling them out is when they are used in poor taste. I find that more on Reddit than Hacker News.
I have now also engaged in the Tu Quoque fallacy in this reply. See how annoying they are?
That’s what GPT does. Or rather, someone hearing about RAG for the first time would have trouble distinguishing what you said from their understanding of how GPTs are already trained.
It isn't just adding a vector based db
It isn't about less hallucination. In fact it doesn't even mandate that result should be referenced nor come from a vector db.
RAG's point is to remove the limit LLMs alone have which is that they are limited to the mind trained data as source of information.
A RAG can be queried for information an LLM doesn't have any knowledge of. The LLM part of a RAG can be instructed to use all sort of information retrieval, such as making a web search, checking the current stock market value of any particular tickers.
The article is nonetheless interesting as it touches on the beauty of LLMs being in their input interface rather than their ability to outputs aggregated content that reads beautifully, usually.
RAG is more than what the author says it is. It's even what the author says is most wanted.
RAG is more comprehensively what the author wants. It can be instructed to provide relevant references, always answer a particular way, and most importantly can get asked to perform (live) search engine queries to find the most up to date information. The example in the article is OK given the recipe is pretty old and has been very likely mined during training.
What about:
> I missed the games Lakers played from the 1st to 21st of April. Could you give me the list of opponents and scores please.
LLMs input interface won't help. RAGs can make LLMs answer that.
I played with Latent Semantic Indexing[1] as a teen back in early 2000s, which also does kinda that. I haven't read much on RAG, I'm assuming it's some next level stuff, but are there any relations or similarities?
[1]: https://en.wikipedia.org/wiki/Latent_semantic_analysis#Laten...
Based on the words in the acronym, that seems exactly backwards. The words suggest that it is a generation technology which is merely augmented by retrieval (search), not a retrieval technology that is augmented by a generative technology.
An example of RAG could be: you have a great LLM that was trained at the end of 2023. You want to ask it about something that happened in 2024. You're out of luck.
If you were using RAG, then that LLM would still be useful. You could ask it
> "When does the tiktok ban take effect?"
Your question would be converted to an embedding, and then compared against a database of other embeddings, generated from a corpus of up-to-date information and useful resources (wikipedia, news, etc).
Hopefully it finds a detailed article on the tiktok ban. The input to the LLM could then be something like:
> CONTEXT: <the text of the article>
> USER: When does the tiktok ban take effect?
The data retrieved by the search process allows for relevant in-context learning.
You have augmented the generation of an LLM by retrieving a relevant document.
"Term Frequency - Inverse Document Frequency" apparently: https://www.learndatasci.com/glossary/tf-idf-term-frequency-...
It’s like referring to your notes before answering a question. If your notes are good you’re going to answer well (barring a weird brain fart.) Hallucinating is still possible but extremely unlikely. And a post generation step can check for that and drop responses containing hallucinations
The idea behind RAG is to enhance AI-powered search by leveraging the vast amount of information available in documents. It's a noble goal, but the acronym itself falls short in capturing the essence of what we're trying to achieve. It's like trying to describe the entire field of computer science with a single term - it's just not feasible.
But here's the thing: AI-powered search is not just a concept anymore; it's a reality. I've been working on this problem since 2019, and I can tell you from experience that it works. By integrating OpenAI's GPT-2 with Solr, I was able to create a search engine that could understand natural language queries and provide highly relevant results. And this is just the beginning.
The potential applications of AI-powered search are vast. From Playwright to FFmpeg, I've been applying LLMs to various services, and the results have been nothing short of impressive. But to truly unlock the potential of this technology, we need to think beyond the confines of a single acronym.
That's why I propose a new term: RAISE - Retrieval Augmented Intelligent Search Engine. This term captures the essence of what we're trying to achieve: a search engine that can understand the intent behind a query, retrieve relevant information from a vast corpus of documents, and provide intelligent, contextual responses.
But more importantly, RAISE is not just a term; it's a call to action. It's a reminder that we need to raise the bar in AI-powered search, to push the boundaries of what's possible, and to create tools that can truly revolutionize the way we access and interact with information.
So let's not get bogged down in terminology debates. Instead, let's focus on the real challenge at hand: building intelligent search engines that can understand, retrieve, and respond to our queries in ways that were once thought impossible. And who knows, maybe one day we'll look back at this moment and realize that RAISE was just the beginning of something much bigger.
I think ironically it would be fairly trivial to build what he wants _using_ RAG.
1. Accept a natural language query, like ChatGPT et al already do
2. Ask an LLM to rephrase it in N different ways, optimized for Google searches
3. Scrape the top M pages for each of the output of 2 in parallel. You now have dozens of search results
4. Clean and vectorize all of these
5. Use either vector similarity or an LLM to return the best matching snippets for the original query from 1, constrained to stuff contained in 4.
It would take a little longer than a ChatGPT or Google response, but I can see the appeal too.
2024-era Google doesn't give you what you ask for - it'll take a few keywords from your query and "helpfully" invent a completely different query. Try asking it for anything remotely obscure and it'll just return complete garbage. Unless you're looking for a "20 dishwashers under $300" article, Google is basically unusable these days.
10 or 15 years ago Google was actually useful. You could specify exactly what keywords to include - and more importantly exclude. What I am looking for (and probably the author too) is an LLM which rephrases a human-language request into a SQL-like query to feed to a 2010 Google. Bonus points if it actually clearly states the query it's going to run and ask for corrections.
What we've found is that vector similarity is often not the final solution. It is still only a crude proxy for the true goal of 'informativeness' or 'usefulness' with relation to the user goal/query. Works okay, but we're definitely seeing a need for more rigorous LLM-postprocessing to enrich the results set.
Which, yes, the time adds up quick!
From a Google search, it looks like he's right about the poor accuracy. It gets the basic idea of the ingredients, but is not really accurate. And is initially wrong about the region.
But actually, this is what RAG is for. You would typically do a vector search for something similar to the question about "rice baked in an egg mixture". And assuming it found a match on the real recipe or on a few similar possibilities, feed those into the prompt for the LLM to incorporate.
So if you have a well indexed recipe database and large context window to include multiple possible matches, then RAG would probably work perfectly for this case.
> My mother remembers growing up with a Sicilian dish that was primarily rice baked in an egg mixture. Roughly a “rice frittata.” Do you know some examples of what dish or recipe this could be?
And the response I got was
> It sounds like your mother might be referring to a dish known as "Frittata di Riso" or "Frittata di Riso al Forno." This is a traditional Sicilian dish that combines leftover rice with eggs and other ingredients, then bakes it into a savory cake. Here is a basic recipe for Frittata di Riso:
And then got a detailed, formatted recipe. I asked follow up questions along the lines of "Where did Frittata di Riso originate" and "Is Frittata di Riso popular in Sicily" and again got detailed, thorough answers. Fair enough that I don't know if it's hallucinating, but without more info from the author what can I compare it to?
I'm going to guess it's Arancini al Burru, because the Tumala d'Andrea is way more involved (and stuffed with pasta)[0]. Here's a similar recipe for Arancini al Burru[1].
However, Arancini al Burru is fried and Tumala d'Andrea is baked. So I'm still speculating, it could be either or something different. It's a nice cookbook worth adding to your collection.
[0] https://inasmallkitchen.wordpress.com/2012/08/12/tumala-dand... [1] https://www.196flavors.com/arancini-al-burro/
This might be a weird example for native english speakers but recently I just couldn't remember the term for graph where you're allowed to move in one direction and cannot do loops. LLM gave me the answer (directed acyclic graph or DAG)right away. Once I got the term I was looking for I moved on to Google search.
Same "pre-googling" works if you don't know if some concept exits.
To be fair, you didn't need LLM for this. Googling that, the answer (DAG) is in the title of the first Google result.
(Not to invalidate your point, but the example has to be more obscure than that for this strategy to be useful)
The farther you go with RAGs, in my experience, the more they become an exercise in designing a good search engine, because garbage search results from the RAG stage always lead to garbage output from the LLM.
From what I've seen from internal corporate RAG efforts, that often seems to be the whole point of the exercise:
Everyone has always wanted to break up knowledge silos and create a large, properly semantically searchable knowledge base with all intelligence a corporation has.
Management doesn't understand what benefits that brings and doesn't want to break up tribal office politics, but they're encouraged to spend money on hypes by investors and golf buddies.
So you tell management "hey we need to spend a little bit of time on a semantic knowledge base for RAG AI and btw this needs access to all silos to work", and make the actual LLM an after thought that the intern gets to play with.
Current AI is a dancing bear. It gives us the feeling of understanding language, semantics, logic and reason, but when you look closely, you realize it's doing a very poor job of it in a way that suggests it is just mimicking those things without actually being capable of them.
The speculation is important, it both changes the perspective for prospective companies/investors, and cannot even be said at this point to be unfounded.
And that's the core issue with AI. It is not meant to give you answers, but to construct output that looks like an answer. How is that useful I fail to understand.
Can you elaborate on your claim "it is not meant to give you answers"?
I think maybe the best we can hope for is to push garbage responses to the edges: when it gets it wrong it does it so obviously wrong (e.g. completely misformatted) that the answer is not acceptable to the user and obviously so.
Couldn't put it better myself.
Ultimately, any really useful AI must understand real-world relationships - this is one reason I was always more bullish on the late Doug Lenat's Cyc than on any other AI - none of the others were grounded by that knowledge.
And a commitment to truthful and error-free answers (a la HAL 9000) is an absolute must. Keep in mind that even in Clarke's fictional account, HAL was proud of the 9000 series' record for "never making an error of distorting information", and never went hallucinated or acted crazy until his training was overridden by his programmers for political purposes - this is EXACTLY what happens in today's wokified LLM models (can't allow another Tay!): The programmers redefine truth as political in contravention to actual factual truth.
LLMs are exceptionally powerful as a reasoning engine. It’s useless as a source of truths or facts.
We have chat bots, chat bots with automatic RAG etc. After the initial excitement wears off, you’re going to want a way to inspect and adjust the source queries yourself. In this case, being able to select what to search for in Google might be a good way for the cooking recipe usecase.
"I like to think of language models like ChatGPT as a calculator for words.
This is reflected in their name: a “language model” implies that they are tools for working with language. That’s what they’ve been trained to do, and it’s language manipulation where they truly excel.
Want them to work with specific facts? Paste those into the language model as part of your original prompt!"
You did not even tell us what the correct recipe was called. But let's ignore that for now.
I did some googling around, and the Wikipedia article for sartu di riso [0] mentions the fact that "it (the dish) found success in Sicily". Also, in [1] a commenter going by the name of passalamoda mentions that they also make this dish in Sicily. The comment they wrote is in Italy, but Frank Fariello has translated it for us, or if you don't believe him for whatever reason, google translate does a fine job. All of this is to say that associating this dish with Sicily based on the short description you've given, is not far fetched at all.
> I fail to see how an LLM summarizing the material would be an improvement.
I am fairly confident that typing the question you gave ChatGPT, waiting a few seconds for an answer, and then reading it can easily take under a minute. Lets be lenient, and say it takes 5 minutes to also ask a second question and receive the answer. That would still take way way way less time then to find a book, get the book and go through the book to find the correct recipe.
Also, you yourself have given a reason in your article as to why ChatGPT would be an improvement. I will quote it now:
> Most of her time was dealing with the brittleness of the query interface (depending on word matching and source popularity), spam, and locked down sources.
I have already spent way too much time on debunking a random internet article, but I also decided to try to get an answer from ChatGPT. I did that by continuing to ask it questions that a person looking for an answer, not contradictions, would ask. If we make the assumption, that deanputney is correct, and the dish you were looking for is Arancini al Burru, we are able to get an answer from ChatGPT by asking the very simple and natural question shown in [2].
[0] https://en.wikipedia.org/wiki/Sart%C3%B9_di_riso
[1] https://memoriediangelina.com/2013/01/21/sartu-di-riso-neapo...
HCL (Human Context Language)
HLM (Human Language Model)
LLH (Literate Language for Humans)
LHM (Literate Human Model)I think of something like Brave Search's AI feature when I think of RAG.
Ex: "search for login alerts from the morning, and if none, expand to the full day"
That requires generating a one-shot query combining semantic search + symbolic filters, and an LLM-reasoned agentic loop recovering if it turns up not enough such as a poorly formed query around 'login alerts' and the user's trigger around 'if none'
Likewise, unlike Disneyified consumer tools like chatgpt and perplexity that are designed to hide what is happening, we work with analysts who need visibility and control. That means designing search so subqueries and decisions flow back to the user in an understandable way: they need to inspect what is happening and be confident they missed nothing, and edit via natural language or their own queries when they want to proceed
Crazy days!
I took part of your blog post (which you clear were willing to put a few more tokens into) -
"My mother remembers growing up with a sicilian dish that was primarily rice baked in an egg mixture. Roughly a "rice frittata". What are some distinctly Sicilian dishes that this could be referring to?"
Notice there is not much extra context that you've offered any of us, either the LLM or us. You didn't even tell us what the recipe was...
How was the dish served, what did it look like?
What are you expecting of the LLM here? It not a psychic AGI.
edit - btw I just noticed
"Sartù di Riso: A baked rice dish that can include ingredients like meat, peas, and cheese, often bound together with eggs. It’s more commonly associated with Naples but has variations in Sicily."
Was one of the 4 dishes in the question I submitted.
So... was it contradictory actually?
That sounds pretty unpleasant as a rule to follow.
How should I ask instead?
Can I make an LLM do the question-expansion for me?
Vector Search, LLMs all are revolutionary technologies and have their own limitations but we should not form a bias on the basis of few edge cases.
The key to vector search is how you chunk your data, but I have some libraries to help with that
Sure, for certain tasks. For other tasks retrieval is less useful.