PDF processing and analysis with open-source tools (2021)
bitsgalore.org
bitsgalore.org
Static compilation means that it will run on most Linux platforms without extra required software.
I believe one aspect of it will remove characters from included fonts that are not used.
It really is quite impressive.
All printing workflows use proprietary formats as input and bind you to one tool or producer. Could Scribus help with that?
- design svg template (eg. in inkscape)
- set placeholders (as svg is just xml -- {%address%} etc)
- replace the placeholders with the actual value & use headless inkscape to produce PDF
$ inkscape rendered.svg --export-pdf=output.pdf
I think there's also a way to do batch processing where you don't need to spin up a new inkscape process for each file (which takes time), but I don't remember how it works anymore.I can produce HTML files from the same XML sources directly: XML ===XSLT===> HTML.
For differences between PDF and HTML versions I have some special tags and attributes in my XML sources.
If I want to change something in the layout, I modify the XSLT script and run the old XML sources through the pipeline again in one go.
There was some up front effort in designing the XML tag system and writing the XSLT scripts. But since my later layout changes were minor, the required tweaks were easy.
For a customer who needed to import large semi-structured legacy Word documents from another company into a database system, I once implemented the following process: The Word documents were converted to a relatively simple homegrown XML format based on the structural elements of the Word document. The resulting XML documents were manually corrected where the structural elements were incorrect. Some special attributes were added inside the XML documents to associate text passages with already existing database keys. When this was finished, an XSLT script was applied that split the large XML file into smaller ones based on this database keys; a human readable prefix, the key and a date went into the file name. These files were converted in bulk to LaTeX and then to PDF. Afterwards, I used a little tool to bulk upload only the fresh PDFs into the correct database entries based on the keys in their filenames.
For one of my side-projects, a C# application, I am using another, object-oriented approach, where I have an abstract base class for reporting and two derived classes, one that outputs HTML and one that outputs LaTeX. The LaTeX output is then fed into lualatex to produce PDFs. You can check out the free Herodotus edition of my (closed-source) Factonaut project at https://www.factonaut.com/ to see it in action.
[1] Using parameter entities for re-usability, such as
<!ENTITY % output_attr SYSTEM "output_attr.ent">
<!ATTLIST foo %output_attr; >
<!ATTLIST bar %output_attr; >
in the DTD referring to an `output_attr.ent` file with the following contents: output (pdfonly|htmlonly|all) "all"
[2] The declaration in the DTD looks like: <!ATTLIST root version (1.0) #REQUIRED >
and the XML must then look like: <root version="1.0"> ... </root>Scribus and inkscape can import pdfs but it would be preferable to use a clean "source" format, pdf is AFAIK meant for output, like lossy compressed images.
[1] https://impagina.org/scribus-scripter-api/ [2] https://berteh.github.io/ScribusGenerator/
That's not a great idea. PDFs are a very "final presentation" format. They don't even really have a strict concept of a block of text. There are even tools which will happily put the separate letters is various places and call it a day. I've done something like that (just inserting short content + signature into a PDF) and do not want to touch PDF editing with a 10ft pole ever again. To preserve sanity, use a different format that actually understands the structure of the content, one step before the PDF is generated.
What you can do is to use a PDF file for all the static content that appears on each page.
The problem with using PDF is that most tools flatten the output, so that makes things a bit difficult.
More on PDF and why it's hard to do analysis [1]. TL;DR
"PDF was never really designed as a data input format,
but rather, it was designed as an output format
giving fine grained control over the resulting document."
By the way, the PDF 1.4 (most widely used version) specification is over 1000 pages![2] To make it even harder, not everybody follow the spec :)[1] https://news.ycombinator.com/item?id=33146364
[2] https://opensource.adobe.com/dc-acrobat-sdk-docs/pdfstandard...
The closest I've come in the past is manually setting xy coordinates to do crops in pdftotext which was pretty time consuming.
I'm sure it wouldn't be too difficult (famous last words) to use annotation objects with structure codes or some such but I'm surprised even after all these years that there isn't something that lets you do this more simply.
The commercial licensing for some of these is a little hairy, too. I've seen rev share clauses which is, uh, hilarious.
Depending on what you're doing, check your licenses!
How about tesseract (https://github.com/tesseract-ocr/tesseract)
There’s even a library for php (https://github.com/thiagoalessio/tesseract-ocr-for-php). Haven’t used it. I did used python Pytesseract & works fairly well.
https://tsdgeos.blogspot.com/2022/03/okular-signature-suppor...
Or whether the signature is correct in all the small details, e.g. used algorithm, the included information.
And then it sometimes depends on the environment. For example, sometimes a digital signature is only considered valid if all the revocation information and all the certificates are included in the PDF.
https://github.com/apache/pdfbox/blob/5b00807463279f1002e245...
Everything about PDF signatures is maddeningly complex.
As for libraries, you can use jsignpdf from the CLI, PyHanko for Python and probably even more that I don't know.