How to Run SQL on PDF Files
rockset.com
rockset.com
Could be wrong.
A calculator:
https://pspdfkit.com/images/blog/2018/how-to-program-a-calcu...
A game of breakout:
https://github.com/osnr/horrifying-pdf-experiments/raw/maste...
The PDF specification is hundreds of pages.
Very cool, but a bit concerning at the same time. I have vivid memories of Flash exploitations and all the malware that came with it. I know we've moved on since then, but I still don't see any reason to allow dynamic content in a page layout format.
I'm wondering if having a PDF that would have data that could self-update if there's an internet connection.
Anyone wanna collaborate on something like this?
I wonder if having a PDF like this that showed you how your stocks of choice are doing would be interesting.
That it's more portable? It's more document-like/printable?
The second one is an empty page. I suppose it needs to do some initialization first in order to show the page.
Thank goodness, though I suppose you could take the position that you already have a full programming language in your document so how could this be worse? Though the idea of distributing node with all my documents seems a bit hairy.
1. brew install pkg-config poppler (on mac)
2. sudo apt-get install poppler-utils (on Debian/Ubuntu)
If you want to extract tables from a pdf, there's Tabula[1], but it isn't automated to run over the whole pdf - you've to do a manual rectangular selection around the table you want to extract.
With Rockset you can avoid ETL when it comes to extracting and manipulating the data. Also, the main value here is that you can join this data with other data sets that are in JSON, CSV, XLS or Parquet formats using SQL to help in analysis.
There are many people who want usable data from such sources. And your service wouldn't be doing any scraping, so you'd probably be OK legally. But IANAL, so do check.
I want a simple library I can use to do this locally for security reasons and because it's a rarity when it happens, but a big problem when I do have it.
I was somehow expecting to have the parser automatically recognize patterns in the PDF and maybe try to name them (and let you rename them), kind of what advanced web scrappers do.
Since its regex only techies would be able to do that in that case they can just write their own script instead.
You are correct that Rockset is doing text extraction for PDF but the main value here is that you can join this data with other data sets that are in JSON, CSV, XLS or Parquet formats using SQL without doing any ETL.
Also, I uploaded my PG&E bill to Rockset and got an empty result set...maybe I'm using it incorrectly.