RAGFlow is an open-source RAG engine based on OCR and document parsing
github.com
github.com
It correctly identifies tables that even paid models like the AWS Textract Document Analysis API fails to – for instance tables with one column which often confuse AWS even if they have a clear header and are labelled "Table" in the text.
I would however love to know broadly what kind of document it was trained on, as my results could be pure luck, hard to say without a proper benchmark
Very nice layout recognition, although I can't quite comment on the RAG performance itself – I think some of the architecture decisions are odd, it mixes a bunch of different PDF parsers for example which will all result in different quality and it's not clear to me which one it defaults to as it seems to be different in different places in the code (the simple parser defaults to pypdf2 which is not a great option)
It seems to be a YOLOv8 fine-tune, I only did a couple tests but results were decent. Another model that is supposed to be fine tuned for borderless is https://huggingface.co/keremberke/yolov8m-table-extraction but I haven't had great results myself with it, but maybe worth a try for you.
I wonder if that will change - eventually.
Discord channels are named in Chinese, though there are English posts.
If instead we had a PDF() class that did what RAGFlow is doing (dealing with all the different trade-offs of the different python PDF engines such as pdfplumber), then we could easily adapt it and improve it, and it can be useful for other projects as well.
Also might be helpful to have a simple example on the README showing how to fetch a document and start querying it. I would try it!
I would like to recommend the FigJam feature in Figma to you, a common tool used by high-tech companies for creating Information Architecture (IA). Our company always adheres to the tool usage standards of high-tech companies.
For guidance on how to use the FigJam feature in Figma, you can refer to this link: https://youtu.be/axDzyLEfYgU?si=V6tqO_tEUKYuLxrL (or search for FigJam on YouTube).
Here are three quick tips on how to efficiently create Information Architecture(Also called System Architecture):
Start with the Sketch method by drawing your product's Information architecture on paper. Communicate with your PM, design, and development team to validate its effectiveness.
Open FigJam in Figma to turn it into an electronic version. This step is crucial as design and development will follow the version in FigJam for collaborative work.
When creating your Information Architecture (IA), first identify the core functionalities of your product, represented by one color. Next, determine what sub-functions each core functionality can be divided into, marked by a second color. Finally, decide what detailed functionalities compose each sub-function, indicated by a third color. Continue in this manner until you complete the entire IA construction.
It’s important to note that perfecting the IA is not a linear process; it will go through multiple iterations and modifications. Every great product undergoes this process. Also, initially, you can focus on creating an IA for just one core functional module (usually the innovative feature with the highest user pain point) without defining the entire scope. In essence, flexibly establishing the IA to achieve company goals is the primary task.
Since "deep document understanding" is not a term of art, I would have just said "OCR and document parsing".
How well does it work? Please include benchmarks. You may be interested in
https://paperswithcode.com/sota/optical-character-recognitio...
https://paperswithcode.com/task/document-layout-analysis
The models seem to be closed source, hosted here: https://huggingface.co/InfiniFlow/deepdoc
[1] https://arxiv.org/pdf/2308.15996v1.pdf [2] https://www.adept.ai/blog/fuyu-8b [3] If you don't want to search for the figures I made a tiny post about it on my weblog: https://weblog.snats.xyz/posts/2024/02/16/
I understand it’s funny to point and laugh “ha ha dumb dudes didn’t think about something female related because they’re dudes”, but it’s worth remembering that not everyone is from the anglosphere.
Also you might find this normal but I cannot imagine my female colleagues being amused at me calling one of them up asking to “consult” about a name solely because she’s a woman.
I'm not even suggesting that this project needs a new name - although I think if they were naming a consumer-facing product or company, someone in marketing would push back on the name almost immediately.
Keep gender out of engineering.
I'd certainly hope it wouldn't be "solely" because she's a woman. There's no people you'd be interested in input from that happen to be women?
A project name is something you'd want to throw around to a few people (ideally with different perspectives) and make sure it conveys what you intend and has the right tone and such.
"On the rag" is a semi-common slang for menstruation, but "rag" has way more meanings than just menstrual pads. IMO, as a marketing term, "RAGflow" sounds about as odd as "Nintendo Wii".
https://web.archive.org/web/20130623080716/http://www.forbes...
http://news.bbc.co.uk/2/hi/technology/4953650.stm
https://www.cbc.ca/news/entertainment/why-wii-ask-nintendo-1...