(Except PDF/A-4, which reintroduces JavaScript for some horrific reason).
(Except PDF/A-4, which reintroduces JavaScript for some horrific reason).
It was pdf.js handling of fonts
Copy paste mostly works fine for me. I only have trouble when it's generated in a weird way (eg. scanned from a paper document then fed through OCR), or has complex formatting (eg. math equations) that have no hope of working correctly in any system. In those cases, I don't see how it's the fault of the PDF format, any more than HTML (or whatever you think is a "real digital document" format) can embed a picture of a scanned document that totally breaks copy-pasting.
(but also math equations have plenty of hope even though they're complex indeed, you can copy&paste some kind of "latex" representation that is sometimes used to ... produce those PDFs)
> whatever you think is a "real digital document" format
whatever supports basic digital interaction we've had available to use for many decades in alternative formats, or whatever doesn't have those rigid pre-digital-paper-based layout limitations where you can't use one of your most popular digital devices - your phone - to read a doc since the phone is smaller than a sheet of paper
> PDF file format (which supports semantic paragraph tags, for example).
These are called newlines and have a pretty widespread support outside of some paper pockets of resistance! You only need some other semantic tags because the format fails at basics
Here is one from Adobe https://www.adobe.com/support/products/enterprise/knowledgec...
Or even better: their annual investor docs a team of professionals has spent time carefully preparing...
like this https://www.adobe.com/pdf-page.html?pdfTarget=aHR0cHM6Ly93d3...
(but don't look at the annual report, that marvel of a public disclosure document not only doesn't copy&paste paragraphs, but has another nice niche use of PDF - you get garbage chars instead of text, rather ironic)
https://www.adobe.com/pdf-page.html?pdfTarget=aHR0cHM6Ly93d3...
[1] https://www.federalreserve.gov/mediacenter/files/FOMCprescon...
The format supports a lot that is not commonly implemented by PDF readers (or PDF producers).
And a good format wouldn't require any ToUnicode maps for simple text in the first place
And poorly supporting a lot without common implementations isn't a defence against the charge of high complexity and bad design, but a reinforcement thereof
(also, no, the first document doesn't work on iOS, I select title and two paragraphs, copy, paste, and I get a single line instead of 3, so a different manifestation of the same common fail of PDFs)
Still, the fact that some PDF processors can make this work shows that the format isn’t broken “by design”.