Tabula – Extract tables from PDF files
github.com
github.com
It's pretty ugly and not cheap, but the data extraction is absolutely magical.
I very rarely feel joy and excitement when using a tool, especially a PDF-related tool, but it saved our dev team at least 100 hours when we first used it. We have it as an automated part of one of our client flows and they happily pay us way more than they should.
It didn't work everytime, but when it did, it was awesome!
https://techcommunity.microsoft.com/t5/excel-blog/announcing...
https://github.com/camelot-dev/camelot https://github.com/camelot-dev/excalibur
[0] - https://docsumo.com/free-tools/extract-tables-from-pdf-image...
What an awesome tool. I had to post process but both were consistent enough in how they were wrong that it was still faster to extract the register maps
"Jan uary"
Is there any reliable way to deal with this?
So does printing the document to a file and Perl.