How to choose your chunk size for summarizing large documents using LLM
vectify.ai
vectify.ai
2. For document division, a bad chunk size can result in the last chunk being much smaller than the previous ones. As a result, the global summary might overemphasize the last chunk and thus bias away from the core semantics of the document. See https://tinyurl.com/vectifyai-summarization for an example.
3. In our blog, we introduce a straightforward method to automatically decide the optimal chunk size for different documents. We find this new strategy significantly reduces bias in the global summary. The implementation can be found at https://github.com/VectifyAI/LargeDocumentSummarization.
Stay tuned at Vectify AI (https://vectify.ai) for more innovative techniques to turbocharge your LLM pipelines!
The only drawback is that, for every corpus of documents, you would have to write custom chunking code that takes into account how big the documents' sections and sub-sections are, and how they're labelled. But this isn't realistically too demanding, and the difference in quality would probably make it absolutely worth it.
1. ada-002 by OpenAI is capable of creating vector embeddings from text fragments up to 8,191 tokens: https://platform.openai.com/docs/guides/embeddings/embedding....
2. Open Source models are usually limited to 512 tokens. Oddly, sometimes people confuse "sequence length" to mean "length of the string". AFAICT, it's always a limit with tokens: https://huggingface.co/spaces/mteb/leaderboard
2. We can estimate the number of tokens in a string to be ~1-5 characters per token. Sometimes, with some data, the average length of a word in English is ~5. So, with 512 characters, the expect token length will be ~100, give or take. Keep in mind it could be larger.
3. If we consider document fragments, proceeding fragments (the ones occurring in the document just before the high ranked fragment) are likely to be important to context. Similarly, we'd want to also include the fragment just after a high ranked fragment. So, that's ~300 tokens we're "pulling in" on a single "match" to a high ranked fragment.
4. Let's assume we use (4) total top hits of similarity. If we only do the before and after fragments of the top hit (which itself is ~100/tokens), then we have (1 * 3 + 3) * 100/tokens = 600 tokens being considered for building the prompt (which will be inferred by the LLM, which has a much longer token length allowed). If we do all of them, then that's (4 * 3) * 100/tokens = 1,200 tokens.
5. However, if we consider a semantic graph of keywords in the document, these keyterms found in the previous documents will be linked to other document fragments in the text. AND, those keyterms can be stored in a non-vector format, in a database. This is similar to doing a summary by the LLM during indexing... Let's say we take the top 5 keywords and pull in high ranked fragments of those. That may give us, let's say 10, more document fragments we can include that are also relevant to the question/ask represented in the user's query (user prompt).
At this point we have ~1,000 tokens from the semantic graph use, from keyterms, and 600-1,200 tokens from related documents, through the vector comparisons. That's about ~1,600-2,200 tokens worth of embeddings we've pulled total. If we combine that with history, or whatever, then we're pushing the "golden" limit of 2,048 tokens for submitting to the LLM.
A lot of people have the idea that longer prompts are better, but it may be that attention is affected greatly in longer prompts, no matter the context window size for LLMs, or the LLM used. As with humans, keeping it on topic is critical to getting a good answer. So, we say 2K tokens is a good goal, given our actual token length will be unknown. At any rate, optimization is the goal here, not including everything but the kitchen sink.
As for "bad chunk" sizing, a good strategy that doesn't involve indexing summaries is to tack short strings onto other strings, when they are found. Sticking them on the previous fragment or the next fragment (if there is one) is a good way of handling them.
Here's an implementation of this strategy for PDFs: https://github.com/FeatureBaseDB/DoctorGPT. An improved version of this code will be pushed this coming week.