Lightweight PDF parser with layout, tables, formulas and bounding boxes
github.com
github.com
Papero seemingly fails to determine any structure here.. for reference, this is the bash script I used to process credit card statements from Scotiabank. While this doesn't generate structured data (which I could certainly do with more work), it generates text files which preserve the layout, making it greppable, which is good enough for my needs right now.
for f in *.pdf; do
f="${f%.pdf}"
if [[ -f "${f}".txt ]]; then
continue
fi
# e.g. "Statement Period Mar 7, 2024 - Apr 4, 2024"
period=$(pdftotext -f 1 -l 1 -nopgbrk -layout -x 331 -y 9 -W 277 -H 15 "${f}.pdf" - | xargs | sed 's/Statement Period //')
startdate=${period%%-*}
enddate=${period#*-}
newname=$(gdate -d "${startdate}" +%F)_$(gdate -d "${enddate}" +%F)
mv "${f}.pdf" ${newname}.pdf
pdftotext -nopgbrk -f 1 -l 1 -x 70 -W 300 -y 276 -H 600 -layout "${newname}.pdf"
pdftotext -nopgbrk -f 3 -l 3 -x 70 -W 300 -y 170 -H 720 -layout "${newname}.pdf" - >> ${newname}.txt
done
Extracting the tabular data here should be pretty straightforward as well, I just haven't needed to do it.LLMs would absolutely eat this kind of task up (writing a script that can turn it into structured data), but I don't have local LLMs set up and don't really wanna send financial records to big AI.
The goal is a lightweight document extraction pipeline that preserves document structure and bounding boxes while exporting to Markdown, JSON, Excel and Word.
I'm also working on structure-aware semantic chunking for RAG, so retrieved chunks can retain their section, page and exact visual location in the PDF.
The project is open source and I'd love feedback on the architecture, extraction quality, and useful use cases.