Show HN: Beyond text splitting – improved file parsing for LLMs
github.com
github.com
IMHO, in most use cases, chunking optimisation strategies will not substantially improve performance. What I think might improve performance is running N search strategies with multiple variations of the search phrase and picking up the best answer. But this is currently expensive and slow.
Having developed a RAG platform over one and a half years ago, I find many of these challenges strikingly familiar.
Your right that chunking is just one piece of this. But without quality chunks you're either going to miss context come query time (bad chunks) or use 100X the tokens (full file context).
> When a user asks a question there is no guarantee that the relevant results can be returned with a single query. Sometimes to answer a question we need to split it into distinct sub-questions, retrieve results for each sub-question, and then answer using the cumulative context.
> For example if a user asks: “How is Web Voyager different from reflection agents”, and we have one document that explains Web Voyager and one that explains reflection agents but no document that compares the two, then we’d likely get better results by retrieving for both “What is Web Voyager” and “What are reflection agents” and combining the retrieved documents than by retrieving based on the user question directly.
> This process of splitting an input into multiple distinct sub-queries is what we refer to as query decomposition. It is also sometimes referred to as sub-query generation.
Eerily similar to Thinking Fast and Slow, and may help explain (when combined with biological and social evolutionary theory) why people have such a strong aversion to System 2 thinking.
The OCR is slow on CPU (working on it), but faster than tesseract (CPU-only) on GPU.
You could probably replace pymupdf, tesseract, and some layout heuristics with this.
Happy to discuss more, feel free to email me (in profile).
The benefit is to let people evaluate surya against the open source and commercial SOTA, improving the integrity and applicability of the benchmark in a business or research setting.
There's a risk: it could make surya's benchmark look less attractive. Also, picking textract to represent the proprietary SOTA might be dicey, since it has competitors (Google cloud ocr, Azure ocr)
Still, ranking surya with doctr, textract, and tesseract would be really nice baseline. As a research user, business user or open source contributor, those are the results I need to quickly understand surya's potential.
Are there any similar projects that are lower level (for those of us not using Python)? Something in Rust that I could call out to, for example?
I see this in the README under the "How is this different from other layout parsers" section.
> Commercial Solutions: Requires sharing your data with a vendor.
But I also see that to use the Semantic Processing example, you have to have an OpenAI API key. Are there any plans to support locally hosted embedding models for this kind of processing?
This implementation bolts on Tesseract which IME is typically not the best available.
[0]: https://en.m.wikipedia.org/wiki/Longest_common_substring
grep -C $n word document
will get you $n lines of context on either side of the matching lines.Using the models underlying a library like this, there's maybe room for fine-tuning as well if you have a set of documents with specific semantic boundaries that current approaches don't capture. (And you spend an hour drawing bounding boxes to make that happen).
Example:
£243,234 would be £234,
Or £243 234
Or £243,234 (correct).
Some cells weren't even detected.
Have you tried this ?
https://camelot-py.readthedocs.io/en/master/
I like Camelot because it gives me back pandas dataframes. I don't want markdown, I can make that from a dataframe if needed
Unitable itself has shockingly good accuracy, although we’re still working on better table detection which sometimes negatively affects results.
I finished watching this video today where the host and guests were discussing challenges in a RAG pipeline, and certainly chunking documents the right way is still very challenging. Video: https://www.youtube.com/watch?v=Y9qn4XGH1TI&ab_channel=Prole... .
I was already scratching my head on how I was going to tackle this challenge... It seems your library is addressing this problem.
Thanks for the good work.