Tabula: Convert PDF Table to CSV
tabula.technology
tabula.technology
Just give me a CLI package that takes a PDF and gives me text file as output.
https://camelot-py.readthedocs.io/en/master/
It's usable from Python or via a CLI.
wholeheartedly agree, and ... give a try to:
pdftotext -layout somePDF.pdf -You just cannot approach it in a fire-and-forget way. It has two modes of operation and various PDF "styles" can respond differently to each mode.
If you have a series of similarly-structured PDFs, try to import them manually (e.g. using IPython), take note of which mode worked better, possibly some adjustments (detection thresholds). Then you can pretty much automate with these collected parameters.
It's great if it works.
AWS Textract is a cloud-only service that you can use either on the console or through the APIs. You cannot run this locally.
Bearer of bad news (sorry), but my experience with tabula has been so-so.
First, installing it is a major PITN.
Second, the output is unpredictable.
Ultimately, I've found that using:
pdftotext -layout somePDF.pdf - | python3 myparser.py
where myparser.py is a 20 lines python script with a couple of regexes and a simple state machine works absolute wonders and can extract relevant data from PDFs even when the data isn't really organized in a table.Also, pdftotext is open source, written in C++, and doesn't require to install a bottomless pit of dependencies like tabula does.
And of course, neither of these things will solve extracting data from PDFs that embed rasterized images or data that is the result of a complex SVG-type rendering.
I believe the real solution to that problem will be: render the PDF to an image at hi-rez and pipe it to some ML-powered process that reverse engineers relevant data out of arbitrary images.
Startup idea?
1 star, won't even try it.
They then wanted a client application but didn't want to build a GUI in java (I assume?) so they grabbed ruby and created a webapp?
That said, it seems like you can run a CLI version using only java and nothing else.