Pdfcpu: A Go PDF Processor
github.com
github.com
Even those with text actually attached / extractable have no structure. "Selecting blocks of text" involves guessing which order the lines go in, depending on their location / distance from other lines.
Compare to having for example "<recipient-address>...</...>" from which you can still generate the printed version.
There's a structure, just not tag-based but position-based ? Ofc, if humans shit around and change it, you're fucked with versionning your masks, but usually they print from form templates themselves.
As I used to say to my colleagues bemoaning this inconvenient analogue bridge: "if you can read it coherently as a human, we can parse it". We have to accept that administrations communicate via geometry and not semantic, and adapt while we also try to convince them to give structured tagging a chance. But they need a critical mass of their documentation pipeline to be machine-read before they even accept to discuss it.
I'm really confused here, it seems we all agree that pdf is a bad format?
We can adapt to scanned documents before all documents are semantically tagged, just like we have to adapt to non standard ascii extensions in non English countries, is my point.
If you create your own PDFs, you can make sure they contain both information about reading order and the mapping from glyphs back to UTF-8 text by creating an accessible PDF (aka a “tagged PDF”)
I think most modern word processors can create such PDFs. MS Word definitely can (https://support.microsoft.com/en-us/office/create-accessible...).
> Compare to having for example "<recipient-address>...</...>" from which you can still generate the printed version.
Generating _a_ printed version is easy; generating _the_ printed version, guaranteeing 100% reproducibility isn’t. To get the exact same layout, you’ll have to guarantee to use the same fonts (difficult, as OSes can update their fonts, possibly tweaking a glyph, a kerning table or anything else that can affect layout) and, basically, never fix bugs in your PDF generation flow.
That’s why many people keep both the structured source data (e.g. in json or xml) and the generated PDF.
I don't think I've seen a tagged PDF in the wild... ever. I'm sure they exist, but I'm doing a lot of stuff with PDFs in the healthcare context and this tech may as well not exist for me. To the point that most apps will support embedding a bad PDF in an HL7 file just to add metadata.
> That’s why many people keep both the structured source data (e.g. in json or xml) and the generated PDF.
They totally should. No dispute.
Oh my sweet summer child.
https://rawgit.com/osnr/horrifying-pdf-experiments/master/br...
Is there a reason you can't have both? Presumably you have structured data at some point, before it's laid out on the page and saved as PDF. Why not just save that alongside the PDFs? You could also serialize it and include it in a PDF metadata field, so it can be extracted from the files even if the database is lost.
I’m hoping pdfcpu does the right thing instead and actually rotates the image in the file.
cpdf -upright in.pdf -o out.pdf
will set the page rotation to zero, counter-rotating the page dimensions and content to compensate, leaving it visually unaltered.(Disclaimer: I wrote it)
I’m guessing they were copy and pasting from PDFs, and unchecked this was breaking the front end system.
Things have changed since 1998, and as hardware has grown so have the requirements for the software that utilizes it.
I wonder if the relative percentages have changed much since 1998.