Breaking up is hard to do: Chunking in RAG applications
stackoverflow.blog
stackoverflow.blog
I’m currently trying to implement chunking by topic using an LLM. It’s much slower, but I hope it will be a huge win in retrieval accuracy. First step is to extract topics from the document by asking the LLM to identify all topics, and then split by sentences and feed each sentence to the LLM to identify the topic. I’m hoping the result will be the original text, split by topics. From there, they can be further chunked if needed. Of course, it could be done by just one shot asking the LLM to summarize each topic in the document, but the more the LLM is relied on to write, the more distortion is introduced. Retaining the original text is the goal and the LLM should just be used for decision making.
Here is the crate I’m working out of. The chunking hasn’t been pushed, but you can see the decision making workflows. https://github.com/ShelbyJenkins/llm_client
I’ve had a lot of success with chunking documents by subtitle. Works especially well for web published documents because each section tends to be fairly short.
I have a chunking library that does something similar to this, and there’s actually quite a few different libraries that have implemented some variant of “semantic chunking”.
I’ve talked through this problem with dozens of startups working on RAG, and everyone has the same problem.
LLMs are arguably not reinventing embeddings if you’re using them to infer the structure of a broad document. Understanding the “external structure” of the document, the internal structure of each segment, and the semantic structure of each chunk is important.
>The problem with most chunking schemes I’ve seen is they’re naive. They don’t care about content; only token count.
I think controlling different inputs depending on context is used in agents. For the moment i haven't seen anything really impressive coming out of agents. Maybe Perplexity style web search, but nothing more.
This is from the first half of 2023 or so; maybe things are more stable now, but looks like the Python implementation is still pre-v1.
Simple concatenation is the obvious way, but I've had better results from using the chunks as anchors for expansion. Use something like Wilson score interval to decide how many high-scoring chunks to pick, then grab the surrounding text for each chunk up to some semantic boundary like period or newline (or arbitrary limit if no boundary is detected).
A tool like Lucene seems far more competent at the task of "find most relevant text fragment" compared to what is realized in a typical vector search application today. I'd also argue that you get more inspectability and control this way. You could even manage the preferred size of the fragments on a per-document basis using an entirely separate heuristic at indexing time.
The semantic capabilities of vector search seem nice, but could you not achieve a similar outcome by using the LLM to project synonymous OR clauses into the FTS query based upon static background material or prior search iteration(s)?
This is also why generally you don't get a vectordb and instead just add a vector index to the DB you are already using
Re:semantic vs synonym, that's really domain dependent. As the number of hits go up, and queries get more interesting, the more vectors get interesting. At the same time, vectors are heavy, so there's also the question of pushing the semantic aspect to the ranker vs the index, but I don't see that discussed much (search & storage vs compute & latency)
Could it make sense to perform dynamic vector lookup over the FTS result set best fragments? This could save a lot of money if you have a massive corpus to index because you'd only be paying to embed things that are being searched for at runtime.
Focusing on just the best fragments could also improve the SnR going into the final vector search phase, especially if the fragment length is managed appropriately for each kind of document. If we are dealing with a method from a codebase, then we might prefer to have an unlimited fragment length. For a 20 megabyte PDF, it could be closer to the size of a tweet.
I really hate it when people throw around acronyms instead of just hitting a few extra keys on their keyboard for clarity
https://learn.microsoft.com/en-us/azure/search/hybrid-search...
I have created a very bare bones implementation as an example at https://github.com/zby/LLMEasyTools/tree/main/examples/agent... (it uses whoosh for indexing) and I am working on a more complete one at https://github.com/zby/answerbot.
the size therefore depends of your content style.
E.g. for a HN discussion, I would go 'paragraph' of each comment.
in a contract, each clause.
in a non-fiction book, maybe each section...
you can also decide to do some kind of reverse adaptive tree: you chunk at sentences level, then compare, if 'close enough', you merge them into a bigger chunk
It sounds like (at the expense of more computation and time) the reverse adaptive tree approach you described would be ideal for those scenarios.
Also, I wouldn’t necessarily use a “sentence” as the lower bound, since that can be something like “Yes.”
[1] - https://docs.unstract.com/editions/cloud_edition#summary-ext...
- Titles matter, a lot: if you add the title of the section at the start of each chunk you will get 10x better embeddings and so more accurate results.
- The size doesn't matter: It depends on the combination of the layout and semantics of the content.
- Avoid garbage in / out: increased context windows don't mean you can put trash inside them. The more good you are at putting relevant information the more precise answers you get. Especially for enterprise-grade solutions, this is so important.
There are good emerging API solutions that implement semantic + layout-based chunking, which in my opinion is the best chunking strategy for PDF / Office files (the widest use case scenario for enterprises).
Cost isn’t as important to us, so we use small chunks and then just pull in the page before and after. If you do this on 20+ matches (since you’re decomposing the query multiple times), you’re very likely finding the content.
Queries can get more expensive but you’re getting a corpus of “great answers” to test against as you refine your approach. Model costs are also plummeting which makes brute forcing it more and more viable.
Also bigger context windows mean a lot more time waiting for an answer. Given the quadratic nature of context windows, we are stuck using transformers in smaller chunks. Other architectures like Mamba may solve that, but even then, increases in context window accuracy are not 1000x.
When you get a query, you then run two semantic search queries: one using the original question and one using a HYDE version of the question. Take those results and run it through cohere’s rerank.
I'm not familiar with HYDE version. I'll check it out. Thanks for the suggestion