PgPDF: Pdf Type and Functions for Postgres
github.com
github.com
1 - https://github.com/turbot/powerpipe 2 - https://hub.powerpipe.io 2 - https://github.com/turbot/steampipe-plugin-openai
It's also pretty easy to write custom plugins once you understand how it's done.
Why would you disclaim something cool?
Certainly interesting, but I think I’m gonna stubbornly stand my descriptivist ground here. We conquered “literally”, and “sneak peak” is soon to fall — I’ll add this to the list!
1 - https://steampipe.io/blog/2023-12-postgres-extensions 2 - https://steampipe.io/blog/2023-12-sqlite-extensions 3 - https://steampipe.io/blog/2023-12-steampipe-export
Last time I searched only https://fdw.dev came up (from Supabase).
pgPDF: The actual PDF parsing is done by poppler.
Poppler is a PDF rendering library based on the xpdf-3.0 code base.
Xpdf is based on XpdfWidget/Qt™, by Glyph & Cog.
XpdfWidget is based on the same proven code used in Glyph & Cog's XpdfViewer library.
The XpdfViewer® library / ActiveX control provides a PDF file viewer component for use in Windows applications.
Quite the rabbit hole!Any licensing complications? Is it cross-platform? XpdfViewer seems to be propriatary and Windows-only.
https://www.xpdfreader.com/download.html
note that the source code is not on github, but is just dumped each version as a tar
https://www.xpdfreader.com/old-versions.html
I once needed to have it for some PDF experiments and I put it on github (this is the newest version; I did _not_ go old versions one by one; I just dumped 2 newest versions)
Functionally it looks useful, but if those kind of 'helpers' catch on there really should be a way to sandbox these 'parser' processes.
The risks of running this code are just way too high without an org level security policy about what access this compromised machine would have.
pgsql -> http -> [ firejail [ deno pdf.js (file:///untrusted.pdf) ] ]"Note that Poppler is licensed under the GPL, not the LGPL, so programs which call Poppler must be licensed under the GPL as well. See the section History and GPL licensing for more information."
See https://gitlab.freedesktop.org/poppler/poppler/-/blob/master...
Pypdf2 and pillow to process images at a high level.
Some clarifications on a few comments I see downstream:
The motivating example was to easily support Full-Text Search (FTS) on PDFs with SQL only (see blog post https://tselai.com/full-text-search-pdf-postgres ). You can treat `pdf` as an alias for `text` and do everything possible.
On the next iteration, I made `pdf` a type (typical varlena object of bytes) to avoid hitting disk all the time. The file is loaded from the disk only once (if it's a valid pdf). One can store the `pdf` type (blob of bytes) as a standard Postgres type. And use that for subsequent calls. Postgres will do it's magic as usual. There is a potential next step of storing the parsed document just to save some time from re-parsing the bytes, but I deemed it a premature optimization.
It's HN's second-chance pool: https://news.ycombinator.com/item?id=11662380
The only functions here all take a filesystem path which your database should definitely not have access to - why would you upload files/store PDFs on a database server!!?
These functions to be able to get the title or modification time of a PDF are also just not that useful.
I don't really understand this question. You can put data where you like. There are no "database servers". These aren't whole toys, take them or leave them. They're made of Lego bricks, and so you can change them.
> COPY naming a file or command is only allowed to database superusers or users who are granted one of the roles pg_read_server_files, pg_write_server_files, or pg_execute_server_program, since it allows reading or writing any file or running a program that the server has privileges to access.
From the documentation:
> Creating a pdf type, by casting either text path or bytea blob.
With the example: SELECT ''::bytea::pdf;
So it's convenient to use the path to test quickly, but you can use anything in PostgreSQL which return (or can be convert) a bytea
Not every object in a database needs to be ready for public consumption. Some of it's there for processing, for certain use cases.
In the second pass I made `pdf` a type.
Some early BLOB implementations had performance problems, which is where this notion that "you shouldn't store files in BLOBs" seems to come from. But modern DBMSes have fast BLOBs, especially if used properly through their streaming API (don't materialize the entire file in memory!).
We have a system in production that stores millions of files in BLOBs, some of them reaching multi-GB sizes, being accessed across the globe (a big enterprise company) by thousands of engineers, and never had performance problems with BLOBs.
In practice though, a PDF for most cases has text-like semantics, so with ::pdf::text you can have all the text-indexes you want.
It seems cleaner to keep this in the service layer and use any PDF parsing library and subsequent schema to store the parsed files.
I wonder what the use case is compared to extracting this information in the programming language and then storing it alongside the PDF in separate table columns?
This also makes interop more difficult when working with indexing databases (i.e. elastic).
This does minimize the amount of client code required to parse pdf's though.