Current architectural best practices for LLM applications
mattboegner.com
mattboegner.com
So far it was really easy to set up the prototype, but the results weren't as great as I had hoped, so I'm excited to see how I could improve it.
Edit: wow, I didn't see this before. LangChain implements one of the featured article's suggestions (HyDE) - https://python.langchain.com/en/latest/modules/chains/index_...
They called it "context injection" but the OpenAI community appears to call it "retrieval-augmented generation".
Edit: I see now that Matt's post is talking all about these ideas. I'm happy that the terminology around this space is solidifying.
(Tangent) I will go to the grave continuing to call it Supabase Clippy even though presumably this prediction from the Supabase blog post became true:
> Today, we're doing our part to support the momentum by releasing “Supabase Clippy” for our docs (and we don't expect this name to last long before the lawyers catch on).
I haven't had a chance to try out hypothetical embedded docs yet, but I expect they only provide a marginal improvement (especially if QAing over proprietary data or information).
I'd love to see any other interesting, more up-to-date resources anyone has found on this topic. I found this recent paper interesting: https://arxiv.org/abs/2304.11062
Can you explain that? I don't follow why it would become less useful
The problem is once there are 10, 20, 30 different-but-similar documents in the vectorstore (like business school case studies), then asking the bot "what are the key takeaways from the airbnb case" grabs a bunch of useless embedded documents to provide as context. Yes, I can tell users how to ask better questions but it's a bad user experience and nobody stick with it or tries to understand why their queries don't work.
I could use hypothetical document embeddings but the problem is a lot of the cases or course notes are proprietary or not publicly available, so I would guess that the hypothetical answers the LLM would come up with won't provide much better context.
This was built with langchain + pinecone.
edit: I think smarter people than I are working on a lot of better ways to do this, but I think one potential solution is to apply metadata to each document when embedding it (e.g., ask the LLM to apply any number of X preset metadata tags) and then, when retrieving context from the vectorstore, filtering the results by those tags.
How would you determine what tags to filter by? Would you also need to rely on the LLM to say "which tags from the collection match this question"?
Sounds like a de-duping problem. Maybe use vector embeddings to find near identical documents and limit them in the context. i.e. maximize the vector distance between your context sources.
Toolformer looks more similar to ChatGPT Plugins (it wouldn't surprise me if ChatGPT Plugins was partly inspired by that paper in fact). It's a way of teaching a model to call tools when it needs to - see also the ReAct pattern.
When you're implementing Q&A on top of a LLM that's not necessarily the right approach. You don't need the model to be able to make searches itself - you need a good strategy for searching for the most appropriate content that fits in the prompt, then sending that to the prompt along with the user's question.
They're both really useful techniques, but I don't see toolformer as a replacement for retrieval augmented generation.
1. Extraction for query filters - https://twitter.com/hwchase17/status/1651617956881924096?s=4...
2. Contextual compression to eek more out of prompt stuffing - https://twitter.com/hwchase17/status/1649428295467905025?s=4...
And then it’s been there’s existing great utility chains for map-reduce, with re-ranking, etc for more ways to apply LLM completions over large documents and/or large sets of documents: 3. https://m.youtube.com/watch?v=f9_BWhCI4Zo
The value the LLM provides in this scenario is simply the language and grammar around knowledge a human provided. It didn’t generate any content whatsoever. The human did that. Similarly a human with a calculator generates numbers but the intent and meaning is provided by the human. The future “LLM all talking to each other” I think doesn’t actually exist, until perhaps multimodal AIs can describe the environment around them. But even then, it’s not an echo chamber because they’re encoding the unfolding of events and time that exists outside the LLM in a way future LLM can encode it.