Local PDF Tools – Powered by WebAssembly
localpdf.tech
localpdf.tech
- https://shreevatsa.net/pdf-pages/ is for extracting pages, inserting blank pages, duplicating or reversing pages, etc.
- https://shreevatsa.net/pdf-unspread/ is for splitting a PDF's "wide" pages (consisting of two-page spreads) in the middle.
- https://shreevatsa.net/mobius-print/ is the earliest of these, and written for a niche use-case: "Möbius printing" of pages, which is printing out an article/paper two-sided in a really interesting order. (I've tried it and love it.)
These don't use WebAssembly, but just use the excellent "pdf-lib" JS library. To keep the file self-contained, I put the whole minified source into a <script> tag at the bottom of the (otherwise hand-written) HTML file.
I'm asking this because an user of my problem validation platform wanted a solution for this[1], because websites requiring document upload have a file size limit and often the compressed file is either above or below the prescribed file size limit thereby loosing out on quality unnecessarily.
[1]'Reduce document file size to specific size' (I have added the link to it on my profile, since it's my own platform).
There's not much to do with fonts except don't embed them unless you need to, and don't have duplicate/overlapping subsets, if you do have these it is very tricky to untangle. I'm not aware of any good tool to do it automatically.
For images, it depends on the format. If your PDF has JPEG (DCTDecode) images then it will have to resample to the JPEG spec, if it's TIFF likewise, you can change the number of colour bits, or you can downsample the DPI which is a simple gs command line option, or you can change the compression scheme within the JPEG itself then replace it. There are so many avenues to approach this that I'm not sure it's something easily achieved while still obtaining a good result across all possible PDFs.
Within a problem domain though, like PDFs that are just pages of scanned images, you could probably iterate and downsample until you hit your target size.
>Within a problem domain though, like PDFs that are just pages of scanned images, you could probably iterate and downsample until you hit your target size.
I presumed any solution for this problem would involve multiple passes through the compression routine to hit the target size. Having to deal with text, font, images separately as you said does make that complex.
Right now, I'm just waiting to see if there's really a need gap for this. It's usually the Govt. websites which has very low limit for document upload size like 100KB, I personally take the image out of .pdf and compress it to minimum jpeg quality; but documents with multiple pages make it tricky.
The first rule of PDFs is to always reproduce at the source if you can, they can be modified/edited but are better considered as an append-only format because of how they are made. There are so many choices that can be made in their construction that are hard to undo later, such as per-character placement instead of per-word or per-line with spaces, each consuming more stream data because of extra overhead in offsets etc (which compresses well as a text stream but still adds up).
Taking out or resampling images like you said is probably the best starting point unless you've found there is a lot of overhead (metadata/unused objects) to trim.
The article that this is based on is here, and a good read. It seems like it's at least non-trivial to get it working, and I'd wonder how the process looks for other compiled binaries, having not tried to do that implementation from scratch. https://dev.to/wcchoi/browser-side-pdf-processing-with-go-an...
https://kc0bfv.github.io/WASM-PDF-Combiner/
I used existing wasm compiles of PDF tools. This use of wasm is pretty awesome to me - I often end up working on very restricted desktop clients with little customization possible, but they always let me run a browser.
Yup - single file operation was a design goal, but I bet there's a better way to get it working than the hack I used.
For Linux and macOS computers that you are allowed to install software on I recommend the pdftk command line tool.
Ubuntu family:
sudo apt install pdftk
macOS with Homebrew: brew install pdftk-javaIf you're on a Mac, the built in Preview tool has had the ability to merge and manipulate PDF documents for years.
I had a bad case of scope creep, so the tool can also extract tables from scanned/image PDFs using OpenCV.js and tesseract OCR wasm build!
I use Tabula under the hood for the cell/row detection and it is really good given the correct mode is selected for the type of table. The modes are stream (find cells by spacing) or lattice (find cells by ruling lines).
The OCR/OpenCV seemed to be fine as well as long as the text isn't too blurry. Here is a GIF of the OCR/OpenCV running on an example Image PDF: https://lh3.googleusercontent.com/-OobUBBtnydg/X6Vn_Ls3juI/A...
[1] https://play.google.com/store/apps/details?id=pdf.shash.com....
I do like this minimal example as a way to get started and see how a very basic PDF is built.
Forgot to say this in my reply, likewise pdftk's uncompress option is my #1 stop for learning about how a PDF is built when we have issues. You can poke around and hex edit for learning too.
It is easy in an uncompressed PDF to comment out a line with % to remove objects that I think are causing issues without messing up the xref offsets. Using a fully fledged tool to edit may have side effects in how things are rewritten which ruins the point of the investigation.
Acrobat Pro has some very useful features too, but not everyone has access to that.
(Ideally I'd like to be able to run such a tool from browser/phone)
Many times I just want to clip white margins from PDFs so that it is easier to view on tablets or phones. Most viewers don't have a way to force the clipping of pages, so when you change page the zoom is lost and suddenly all the content is squished to center.
Last time I found cli programs to do it aprox. five years ago it was really difficult to find good tools to edit PDFs like that.
It's actually not trivial task, as sometimes pages have different margins, e.g. odd and even pages has different margins on folding side of page.
For an ordinary user, the only 'easy' way a user can verify the claimed behaviour is to literally go offline.
Browsers do not currently have badge to verify that the app is not sending any data. I'm thinking how we were brought to trust the padlock icon browsers display for TLS supporting sites.
I just did this and it worked.
I built a tool that required to count the number of pages of a PDF (ca. 2014-2015). At the time server side counting was the 'sure' way in my brief research.