OpenPDF 1.3.0
github.com
github.com
something in C, or anything that compiles through LLVM would probably be a better choice i think.
[0] https://pdfium.googlesource.com/pdfium/+/refs/heads/master/d...
[1] https://pdfium.googlesource.com/pdfium/+/refs/heads/master/p...
It's not packaged as far as I can tell separately from chromium and chromium doesn't ship it as a reusable library. Not sure if it has a stable interface, probably not.
It doesn't have any frontend other than chromium.
Their docs say
"The public/ directory contains header files for the APIs available for use by embedders of PDFium. We endeavor to keep these as stable as possible."
The rest are valid points though!
That sounds more like best-effort.
I feel that the name is very unfortunately chosen - it's too generic, should have had a "J" in it at least to denote that it is a Java library. JOpenPDF, perhaps.
I’ve never seen it work though.
https://www.iso.org/obp/ui/#iso:std:iso:19444:-1:ed-1:v1:en
Though, it was mostly an exercise in frustration. That said, IRS forms used to work great with this.
However, long story short, Evince works most of the time.
Isn't one of the advantages of Java exactly being multi-platform?
You target the JVM. So Java libraries are awkward to call from e.g. C++ or Python.
* https://www.graalvm.org/docs/reference-manual/aot-compilatio...
You only need to pass `--shared`. (You can also define multiple exports via public static ... and some annotations https://github.com/oracle/graal/blob/master/substratevm/READ...) of course it's a little bit akward to first create a "JVM" library and then call that library from native code, but it's not impossible.
E.g. I recently needed to convert raw OCR files to PDF -- place TIFs on each page without recompressing, place invisible text on top, add bookmarks, save. I ended up stringing together three python libraries (img2pdf for efficient image combining, reportlab to place text, PyPDF2 to merge the two), which was both clunky and slow. Second pass was scripting command line tools (tiffcp + tiff2pdf + reportlab + pdftk), which was still clunky but at least faster.
It would be awesome if there was one solid, fast implementation that lots of language ecosystems could wrap and contribute back to. I haven't looked around much outside of Python -- maybe that's OpenPDF with Graal, or PDFium, or something else mentioned in the sibling comments? But there are so many usecases for PDF, I wouldn't be surprised to find that those also fail to cover the whole field.
Y'know what I can imagine working is a package ecosystem that uses a common in-memory data structure, so you have one base package that does the reading and writing and low-level manipulation, and someone else can come along and write the equivalent of img2pdf or reportlab as separate packages that compose without any efficiency loss. That could work in any language but it feels like the kind of thing Rust is doing really well these days. It might just take some good branding around "hey, here's a low-level FooPDF data structure that is easy to write high-level libraries against -- make sure to mention in your docs that you're compatible with FooPDF." Sort of an Apache Arrow approach.
Similarly, I would like to see more attention given to PDF/A and high quality reference implementations and libraries targeting it. It solves many of the most common complaints here on HN about PDF: no audio / video, no javascript, no XML forms, all fonts have to be embedded, etc. - and it's an ISO standard.
We really need solid support for a document format that has readers that aren't constantly getting CVE'd from supporting too much of a bloated specification.
Pdf/a needs a LOT more publicity.
Currently I'm using a 10 year old paid version of Acrobat Pro because I'm unwilling to spend another $500 to get a modern copy simply to compress images. I use it for scanning and archiving old computer manuals. In many cases jbig on 1bpp images is good enough; other times I want to use one of the more sophisticated hybrid schemes so I can have B&W or color images interspersed on the page with 1bpp text. Acrobat does a pretty good job of auto-deskewing images and doing analysis to figure out which compression type to use in different sections of the same page.
https://github.com/4lex4/scantailor-advanced
I find the autodeskewing algorithm to work well, but it allows hand adjustment as well, which I like. As I've gotten better as using it, I've been able to get the size of my scanned documents down considerably by cleaning up the scans. This includes some old manuals.
As far as the pdf encoding itself, I use both mutool, from mupdf, and qpdf. I just checked and it looks like while they both compress their streams, it may not have the same flexibility with Acrobat Pro. For me, I'll decompress, edit, and recompress streams on the files and that's been fine for my use.
That said, if someone knows of a better tool for compressing streams in a PDF, I'd be interested to hear about it as well.
ps2pdf LARGE.pdf SMALL.pdf
I believe it compresses all of the images independently, instead of converting the entire thing to an image as I believe Imagick does (I could be wrong though). This tool doesn't appear to have any intelligent auto-deskewing though. (On a slightly different note, it's questionable whether you want to do this for archiving as it'll undoubtedly be a lossy operation.)The reason I use this: Many websites for conferences have upload limits at around 10MB, a size you can easily reach with a handful of images from a modern device in your paper.
[1] https://www.shellhacks.com/linux-compress-pdf-reduce-pdf-siz...
ps2pdf does work on Windows though (or maybe it's pstopdf?), and the Linux version works directly in WSL.
As the PS format was designed for printing [2], my guess is that it selects a DPI value suitable for printing. I can't seem to find the source code to back this up though. (Ghostscript has the simple web front-end that doesn't lend itself very well for viewing code in browser [3].)
[1] https://www.biu.ac.il/os_site/documentation/gs/Ps2pdf.htm
They have an enterprise PDF compression toolkit which clearly supports all the advanced compression schemes of PDF, but it is too rich for me: prices are not given, just a button to ask for a quote. There is a personal PDF tool for $129 but they don't make clear if it also supports their top quality PDF compression.
PDF.js seems too slow for my purposes, though I haven't tested very thoroughly yet. I'm hoping with wasm Skia now available, there might be some other options coming.
I recently started playing around with wasm Skia/CanvasKit but I have a hard time wrapping my head around what exactly it is supposed to be used for, as well as Skia in general.
For example, Skia allows text rendering and I could naively assume that if I want to build some high performance 2D UI that Rendering text with Skia would work reasonably well.
But at what level would text interactivity like selecting text happen?
I’m afraid this is a very generic question.
Are you familiar with any pdf or text related project that uses Skia and could help me better understand what role Skia plays?
OpenPDF [1]:
document.add(new Paragraph("Hello World"));
PDFBox [2]: contentStream.beginText();
contentStream.setFont( font, 12 );
contentStream.moveTextPositionByAmount( 100, 700 );
contentStream.drawString( "Hello World" );
contentStream.endText();
[1]: https://github.com/LibrePDF/OpenPDF/blob/master/pdf-toolbox/...[2]: https://pdfbox.apache.org/1.8/cookbook/documentcreation.html
Within Polar (https://getpolarized.io) we support pdf.js which is rather nice but ONLY supports viewing of the PDF and text extraction.
You can't create NEW PDFs.
Our plan is to use something like OpenPDF to get the best of both worlds. Editing PDFs for doing things like exporting could be done on the server and the rest done on the client.
If not, is there any that are recommended? PDFTK is really showing its age in an application we have (its slow, and often unreliable in large batch jobs)
Other question: Is it possible to build a document format similar to PDF around SVG?