LlamaCloud and LlamaParse
blog.llamaindex.ai
blog.llamaindex.ai
For character extraction, LlamaParse use a mixture of OCR / character extraction from the PDF (it's the only parser I'm aware of that address some of the buggy PDF font issues, check the 'text' mode to see raw document before reconstruction), use a mixture of heuristic and Machine learning models to reconstruct the document.
Once plug with a Recursive retrieval strategy, allow you to get Sota result on question answering over complexe text (see notebook: https://github.com/run-llama/llama_parse/blob/main/examples/...).
AMA
You could also try https://github.com/VikParuchuri/marker for general PDF parsing (I'm also the author) - it seems like you're more focused on tables.
1. The comparison with the open source pdf libraries is rather strange. They are self contained libraries. Not ML augmented services. Does this mean you plan to open source the underlying technology and offer it under the same licenses as pypdf or pymupdf?
2. How does this compare to AWS Kendra? I have one of the bigger deployments out there and am looking for alternatives.
3. How fast is the pdf extraction? In terms of microseconds given we pay for execution time.
(1) The "baseline" comparison was to PyPDF + Naive RAG. For the LlamaParse evaluation, you appear to have used a different RAG pipeline, called "recursive retrieval." Why not use the same pipeline to demonstrate the improvement from LlamaParse? Can you share the code to your evaluation for LlamaParse?
(2) I ran the benchmark for the PyPDF + Naive RAG solution, directly copying the code on the linked LlamaIndex repo [1]
I got very different numbers: mean_correctness_score 3.941 mean_relevancy_score 0.826 mean_faithfulness_score 0.980
You reported: mean_correctness_score 3.874 mean_relevancy_score 0.844 mean_faithfulness_score 0.667
Notably, the faithfulness score I measured for the baseline solution was actually higher than that reported for your proprietary LlamaParse based solution.
[1] https://github.com/run-llama/llama-hub/tree/main/llama_hub/l...
Thanks for running through the benchmark! Just to clarify some things: (1) The idea is that LlamaParse's markdown representation lends itself to the rest of LlamaIndex advanced indexing/retrieval abstractions. Recursive retrieval is a fancy retrieval method designed to model documents with embedded objects, but depends on good PDF parsing. Naive PyPDF parsing can't be used with recursive retrieval. Our goal is to demonstrate the e2e RAG capabilities of LlamaParse + advanced retrieval vs. what you can build with a naive PDF parser.
(2). Since we use LLM-based evals, your correctness and relevancy metric look to be consistent and within margin of error (and lower than our llamaparse metrics). The faithfulness score seems way off though and quite high from your side, so not sure what's going on there. maybe hop in our discord and share the results in our channel?
This is my problem with projects that start off as open source and become famous because of their community contributions and attention, then the project leaders get that sweet VC money (or not) and make something proprietary.
We've seen it with Langchain and several other "fake open source" projects.
Why shouldn’t they make money? LI is a fantastic way to do RAG.
It could still be licensed in a restricted way, but keeping secret how it works is unfortunate - it breaks the chain of learning that is happening across the open ecosystem and, if the technique is any good, all it does is force open models to build an actually open equivalent so that further progress can be made (and if it's not really any good then it's snake oil, which is worse). Even if it's great it essentially becomes a dead end for the people who actually need and want an open model ecosystem.
The problem is not that it's a hassle to setup and maintain, but that discovery and network effects are a lot smaller when self-hosting. Being on Twitter and Medium gives you a lot more readers from the get-go, just because of the built-in discovery.
I still self-host my own stuff, but I would lie if I said I didn't understand why people use Twitter and Medium for hosting their content.
https://github.com/Unstructured-IO/unstructured/blob/d11c70c...
1. The sign up with email just endlessly redirected, click link in email, ask to sign up with email, put in email, click link in email, etc.
2. Fine, I'll sign in with Google.
3. A PDF parser? Seriously that's what all this fuss is about? There are so many options already out there, PDFBox, iText, Unstructured, PyPDF, PDF.js, PdfMiner not to mention extraction services available from the hyperscalers. Super confused why anyone needs this.
Specific to the parser, they do show where tools like those you mentioned fail and their LLM based parser captures the full data the aforementioned miss.
It's easy to cherry pick a PDF for marketing purposes and claim you're better. I didn't miss it, I just don't believe marketing announcements at face value. I tried their parser on a PDF with a bit of complex formatting like multiple columns, tables and a couple images and it choked, spitting out one big markdown header with jumbled text. Not impressed.
Planning broader releases in the future for sure.
Great work by llamaindex team. Also feel free to try https://github.com/nlmatics/llmsherpa which takes into account some of the things I mentioned.
In LlamaIndex for example, there are a a few markdown-specific classes that work well with this.
You can find an example over in the repo -- https://github.com/run-llama/llama_parse/blob/main/examples/...
I'm sure that for LI, having it as part of a workflow and history to retrieve for RAG makes it easier for users, but why reinvent the wheel?
Unless I missed something
Did you intend to rule out OpenAI from consideration?
You mentioned hardware being a constraint, but that doesn't tell me why you specifically wanted to find an alternative to OpenAI.
Also I would like to pay for an equivalent alternative that is less censored, like ChatGPT had a bug one day that it refused to tell me how to force a type cast in TypeScript, it showed me a moderation error. So I want an AI that is targeted for adults and not children in some religious school in USA.
at this time, your only option is local models. if you don't have the hardware to run them yourself, there are plenty of hosts - poe/perplexity/together etc.
llama3 is (hopefully) coming soon, and if it has improved as much as llama2 improved over llama1, and provides at least 16k baseline context size, it will be in between gpt3.5 and gpt4 in terms of quality, which is mostly enough.
Others include Runpod, Replicate, probably others.
Disclaimer: I'm a software engineer at AI21.
It's the trick where a user asks you a question: "Who worked on the billing UI refresh last year?" - and you turn that question into a search against a bunch of private documents, find the top matches, copy them into a big prompt to an LLM and ask it to use that data to answer the user's question.
There's a HUGE amount of depth to building this well - it's one of the most actively explored parts of LLM/generative-AI at the moment, because being able to ask human-language questions of large private datasets is incredibly useful.
Explained by gpt itself as if you were a teddy bear.
----
Okay little teddybears, let me explain what retrieval augmented generation is in a way you can understand!
You see, sometimes when big AI models like Claude want to talk about something, they may not know all the facts. But they have a friend named the knowledge base who knows lots of information!
When Claude wants to talk about something new, he first asks the knowledge base "What do you know about X?". The knowledge base looks through all its facts and finds the most helpful ones. Then it shares them with Claude so he has more context before talking.
This process of Claude asking the knowledge base for facts is called retrieval augmented generation. It helps Claude sound smarter and avoid mistakes, because he has extra information from his knowledgeable friend the knowledge base.
The next time Claude wants to chat with you teddybears, he will be even better prepared with facts from the knowledge base to have an interesting conversation!
1. Build janky open source code base
2. Sell compute to run it
3. Build features that create compute lock in (vercel is a master at this)
Spend seed round investments on building a solid software but not building an income stream that can satisfy investors, thus not receiving any new funding and let the company die.
Maybe there is a class of developer out there that doesn't get spooked by that but it definitely has created an adversarial place for Vercel in my mind. I feel like I need to be careful when touching anything Vercel have touched so that I don't fall into a trap.
I'm calling this situation Fauxpen Source. The recent moves definitely feel anticompetitive or at least trying to force you into using their products
I'm migrating to vite+vike (next/nuxt like experience for any framework)
[1] https://nextjs.org/docs/app/building-your-application/deploy...
Beyond the runtime limitations, it is poorly designed and requires you to effectively write a router when the rest of the system has automatic routing assembly
40 years after PostScript and this is still a problem that one needs to throw AI at. I feel the software development and human-computer interaction took a wrong turn along the way. What happened to the semantic web?
We still have 'the web'. PDFs are something different and separate.