In my experience Azure’s Form Recognizer (now called “Document Intelligence”) is the best (cheapest/most accurate) PDF parser for tabular data.
If I were working on this problem in 2024, I’d use Azure to pre-process all docs into something machine parsable, and then use an LLM to transform/structure the processed content into my specific use-case.
For RAG, I’d treat the problem like traditional search (multiple indices, preprocess content, scoring, etc.).
Make the easy things easy, and the hard things possible.