Tabula: Extract Tables from PDFs
tabula.technology
tabula.technology
I ended up using pdfquery package in python which heavily utilized PDFMiner under the covers.
Besides ABBYY soft (which is proprietary, licensed), does anyone have other recommendations?
I'd even pay money to get somethig that works well.
Those of us trying to extract the data bound up in these PDF's do advocate to get access to the original data, but we have to deal with what we have today.
My school district (What a mess) publishes images (Horrible bad images) of all the school notes including all financial information and spreadsheets. I had to one night type in for 4 hours manually the years budget just to check on our spending per student. It was $5,400 the lowest in our state.
If I need more detailed formatting information, I use "pdftohtml -xml -fullfontname" and process the resulting xml.
Unfortunately it looks like the developers of JPedal decided to discontinue the LGPL version and focus on the proprietary version, so it's unmaintained unless someone else picks up development.
We use JPedal for rendering pages as images. For parsing, we use Apache PDFBox. In the near future, we plan to render the PDFs client side with Mozilla's PDF.js
Moving to PDFBox 2.0 is also on our roadmap. But the text extraction API in 2.0 has changed a lot too, so porting our engine would require quite a bit of effort.
Friendly reminder: we're an MIT-licensed open source project, and we're always open to contributions!
> Tabula only works on text-based PDFs, not scanned documents. If you can click-and-drag to select text in your table in a PDF viewer (even if the output is disorganized trash), then your PDF is text-based and Tabula should work.