PDF Is the World's Most Important File Format
motherboard.vice.com
motherboard.vice.com
But the ability to collect papers, books, documents, etc. all in a single format that I can read on any device, and mark up with highlights and notes, has been a game-changer.
Yes it's a lowest-common-denominator format. And it's designed for human reading and manual office tasks, not computer processing of data. But it works. And it's supported everywhere.
Doesn't matter if I'm on the Apple or Google or Microsoft or Adobe stack. Doesn't matter if the PDF is 20 years old. It just works.
Just look at HTML. Would you rather have it be impossible to extend beyond text, images and stylesheets?
I get your point, though.
AFAIK DJVU is a much simpler format primarily for scanned images, not vector graphics (which includes LaTeX-generated documents, mind you.) Could anyone with more knowledge about DJVU comment?
I haven't seen that happen yet though...
A "multi-view" format that could be viewed as full page layout (both scrolling and paged), small screen linear, audio-only linear, or machine-readable (full, deterministic text sequence, tabular data, alt-text and tags for images and charts, etc.) All pages, images, fonts, etc., would remain encapsulated within a single file.
The closest I've seen is the latest ebook format, but that doesn't have anything close to the power of PDF to lay out a proper textbook or magazine article page.
If you want something to print beautifully, you take the time to put the figures and pictures exactly where you think they fit. If you want something to be read with arbitrary text sizes, you hope they'll show up somewhere near the relevant text. Treating a book as an ebook, or vice versa, will always be a compromise.
What are you using for that? In my experience annotation support is very spotty, not user-friendly and not cross-platform, so I rarely use it for collaboration. Even fill-in forms one gets for reimbursements often seem badly done with too small or too large boxes or missing boxes. Thus, pdf seems fine for read-only, but making changes afterwards is not a good experience.
For instance, check out the Canadian passport simplified renewal form [1]. The upper-right corner on the first page is "$FORM$054(06-2018)$V$1.4$CS$0$C$0" on the Apple stack and a proper QR code which changes as you fill in the application in Reader. The big blue "Read Instructions" buttons don't work on the Apple stack either.
It may be important, but it's a waking nightmare of a spec.
[0] http://mariomalwareanalysis.blogspot.com/2012/02/how-to-embe...
[1] https://www.canada.ca/content/dam/ircc/migration/ircc/englis...
Adding friends on WeChat, making payments to vendors, getting discounts, installing an app, etc
Then, Adobe saw an opportunity make some enterprise money and introduced XFA (XML Forms Architecture). The XFA spec is associated with the PDF spec, and you can put XFA content inside of a PDF, but the XFA spec is actually larger than the PDF spec. It's utterly insane.
I know that Adobe of course supports XFA, and there are various enterprisey things that support it to varying degrees, but I don't think it's well-supported by anyone outside of Adobe's implementation. Not only is it huge and complicated, but it also requires pulling in a JavaScript interpreter, which is a big ask for a feature that only exists because at one time Adobe thought they could turn it into another revenue stream.
It's noteworthy that the PDF 2.0 spec specifically says that XFA is not only deprecated, but that any PDF which contains XFA content is considered out of spec for PDF 2.0. 2.0 goes back to AcroForms for all of that stuff and ditches XFA entirely. Likewise for JavaScript in general. Anything with JavaScript dependencies in PDF 1.7 is verboten in 2.0.
In general, if you get stuck with a PDF with XFA content (not AcroForms), your best bet is to just use Adobe Reader to fill it out. Hopefully PDF 2.0 will take over the world eventually and everybody will be 100% back on AcroForms.
[1] https://theblog.adobe.com/taking-documents-to-the-next-level...
I won't believe it until I see hard evidence. Acrobat and Reader have supported JavaScript for something like two decades. (I wrote a bunch of JavaScript multimedia APIs when I worked there around 2002.)
Consider the large number of interactive PDFs on the IRS website that use JavaScript for form calculations. I don't think Adobe is highly motivated to break all of those.
I understand this is a problem as it is one extremely important features of PDFs, but from a consumer perspective (never worked in an big office) I do not think I have ever seen a PDF form.
Maybe the real success of PDFs is that it managed to hide all the inner complexity of the format and missing features from the common use cases.
PDF is not the "lowest common denominator".
That would be the format that can be most easily converted to other formats.
With text and images, I can easily create PDFs and documents in myriad other formats.
Alas, PDFs do not convert as easily to other formats. To this day, no one has a PDF-to-text converter that is 100% reliable in preserving the proper line lengths. Not even a company with Google's resources.
But we'll never have that (not in the sense we have PDF.) Either it's an image or an application, no one really understands how to handle anything in between.
Then you have abominations like embedded flash...
As a standalone file representing the format of a book, PDFs are a good format. But then PDFs can unfortunately (sometimes) store much more, and then can be a security minefield.
That said, my personal favorite for a complete screw you to the format comes from Texas and form 2382:
https://hhs.texas.gov/laws-regulations/forms/2000-2999/form-...
If you decompress it, you'll find that they encoded an HTML webpage into the PDF and, as far as I can tell, this can only be viewed on Windows with Acrobat. It's almost like Texas doesn't want people to apply for medicaid funds.
https://en.wikipedia.org/wiki/XFA
Here is a plain Acroform version :-)
https://send.firefox.com/download/3d2da3ccd62d4ce9/#CJRCa5XO...
If you see a way to, make them pay for it; there is money to be made.
If you asked Adobe, they would say this is not possible. :-)
(Note, breakout game only works in Chrome's PDF reader)
https://www.iso.org/obp/ui/#iso:std:iso:19444:-1:ed-1:v1:en
That said, the workflow was pretty terrible. Basically, I had to use pdftk to output from the pdf what fields were writable. Some forms like from the IRS, were pretty good at labeling things. Most other forms were terrible and there was a lot of guessing and checking. As such, this doesn't exactly answer your question, but more to comment that there was sort of a way, kind of.
I would rather have a plain ole html file with formatting flags than a PDF book.
If your pdf book does not have it, there's some readers who can add it. e.g. https://helpx.adobe.com/acrobat/using/creating-pdf-indexes.h...
It also doesn’t render well on different resolutions.
PDF is a perfect case study how inferior solutions can become standards.
Its literally a replacement for paper documents, except you can search and view them on a computer and still hit print and get a great paper reproduction.
Word processing/editing has so much more going on, PDF actually removes features from postscript for security purposes, nobody wants to be sending word docs around just for viewing.
It a better situation than OCRing a paper doc though.
I have policy and legal documents from 2005 in PDF/A that can be rendered identically in 2019, and likely in 2105. That isn't the case for HTML, Word or almost any non-plaintext format. If for no other reason than the US Federal Courts require use of PDF, the format will exist and be somewhat vibrant for many decades to come.
I can wholeheartedly will agree that Adobe Reader sucks, but the format solves lots of problems that are difficult to solve otherwise.
A heavily restricted subset of HTML to replace PDF as the 'archival' format would make the world so much better of a place.
Will HTML formatting on a 4K display will look the same as an 800x600 monitor?
Will all of the ancient display elements display the same? Will IE4 specific artifacts display?
Search works fine in PDF. Mobile is not optimal, but platforms optimized for mobile require different design considerations. Few webpages look or function identically on mobile.
The only use-case I can conceive for perfectly-reproducable layout if you are not a print publisher is in fields where it is convention to reference text by page+paragraph number; in those cases, the page number is actually semantic information so could quite trivially be encoded in the content of the document to maintain that referent.
> Search works fine in PDF.
Search works fine in PDFs that don't use hyphenation or which properly implement it, and when they don't it breaks silently in ways that are potentially disastrous.
> Mobile is not optimal, but platforms optimized for mobile require different design considerations.
And my point about mobile was you don't need to optimize for it.
>Few webpages look or function identically on mobile.
Again, it doesn't need to look identical. That is a falce requirement. There is no value in that. The only web pages that don't function on mobile are ones which have been optimized for desktop. We're not talking about general-purpose web pages here, we're talking about textual documents.
This is not some hypothetical scenario, BTW. The UK has been using HTML over PDF for public-facing documents for a couple years now, it seems to be working out for them[1].
1. https://gds.blog.gov.uk/2018/07/16/why-gov-uk-content-should...
Although it's rare to find them nowadays, this is false for pages with no CSS at all. By no CSS I mean not even the infamous proprietary viewport meta tag... which is a posteriori being made CSS. Access for instance the basically unstyled page [1] with a $1000+ smartphone of our day... and you're likely to find unreadably small font.
Now, you could argue [1] functions on mobile, but let's agree it's stretching the meaning of that term. But that's not the main point: HTML elements come and go (for instance MENU; find others in [2]), so it's clear that archival reliability is not a big priority.
We all agree that ideally the source should always be made available (mandated if tax payers' money is involved if you ask me), but that doesn't invalidate the value of a universal presentational format.
[1] http://www.qrg.northwestern.edu/papers/files/simhobby-local....
This is not true. Ask a Lawyer, Accountant or a Journalist wading through govt pdfs.
The reasons cases and investigations take forever is the time spent manually reassembling related data residing in different orgs. The same data across all orgs will show up as 25 differently structured tables. This is changing slowly but putting public tabular data in pdfs has probably cost the economy billions.
Or you could says it has created a ton of jobs :)
Yes, I have had to deal with it on a development slant as well. The open source tools are rare.
Paper records are the reason systems went to databases to store information.
PDF is NOT the swiss army knife of IT-data and never will be.
Some governments, e.g. Dutch government, mandate that PDFs be tagged PDFs.
Here's computerphile video that explains it
back then there was only 1 format in the world .doc
PDF became standard because it was a SUPERIOR solution for the use cases that most people wanted.
A better solution, perhaps, for preserving structure, would be TeX, markdown or Org files, but PDF has the advantage of not having to be compiled and being ready for presentation/consumption (possessing platform invariance).
And then they added JavaScript to it:
http://www.adobe.com/devnet/acrobat/javascript.html
The small subset that displays text and images is fine, but PDF is a nightmare.I just opened a random PDF on my computer in a text editor and it starts off with "xÕ\ko‹∆˝Œ_1@ÉbD 9|ÌËáƒZçÌDJÇ¢)UZ[nıÚÆÏƒA˛Pˇeœπ"
Its not like you can read text without the right program, its still binary, there just happens to be a mostly agreed upon standard and a lot of programs that can decode that standard and render to screen.
PDF viewers are built into most browsers now and allows rich page perfect print ready results.
If I had it my way we would settle on a binary data serialization format, I don't care MsgPack, Protobuf, heck maybe even sqlite and then everyone could have a viewer to snoop around in it. You would still have to understand whats being encoded but you could always "view source" so to speak.
Same for text. Text can be pure text,but for space saving, it usually will be deflate-compressed.
Hmm. Maybe your charset is Windows-1253, and unassigned characters are being supplanted with the equivalent codepoints from Windows-1252?
- While it's fairly easy to read and write plain text, it's also fairly easy to inadvertently introduce unintended artifacts in the process.
- The more frequently a file gets passed around and read and written to, the more likely mojibake[1] will get introduced. This concern rises exponentially when you move to non-US audiences and introduce local-specific encodings. File storage settings, client operating system settings, server configuration settings, database settings, programming languages that touch it along the way. All of them introduce assumptions along the way of a file's encoding, and many failure cases can be subtle and easily go unnoticed at a glance while causing some irreversible damage to downstream recipients.
- Even if you solve for the encoding, you still have structural issues with tabular data. Different parsers treat escaping and quoting policies differently. This can result in data shifts as things get mis-parsed, data corruptions if literal values get interpreted as escape characters or vis versa, etc.
For preserving data, generic plain text tends to get worse and worse over time because it's such a non-opinionated format and even if you document the specifics on econdings and parsing details it's easy for those to get lost over time as things exchange hands or for intermediaries to corrupt the plaintext because they relied on defaults instead of the documented parsing details.
For better or worse, PDF tends to solve the preservation issue while introducing potential barriers on the parsing/processing side.
These days, those that are promoting the idea of plain text as a long term archive format are assuming UTF-8 by default.
You can see the text code by doing something like this:
pdftk input.pdf output input_uncompressed.pdf uncompress
You can also edit it in that state, in an editor that preserves binary content, but there's a hard-coded offset table at the end, so if you change the length of something, that needs updating (very fiddly to attempt by hand, but automatable, and some pdf tools automatically fix broken offset tables).Maybe you mean ASCII, which is indeed simple and also not useful for several billion people. (I live in one of the many countries that cannot use ASCII -- the American Standard Code for Information Interchange -- because don't speak American here.)
Or maybe you mean Unicode which is an extremely long spec and absolutely cannot be handled by a first year programming student.
It's not uncommon to print things from Chrome/Firefox and have the margins/cropping be wrong. Then again you're using a web browser to print a pdf file, you get what you deserve
Firefox I think still uses https://github.com/mozilla/pdf.js which was not bad last time I used it, but it all javascript so performance isn't up to par with native. Also printing is not great and I don't think they have implemented a SVG backend yet for better printing. On the plus side you can embed it in your web app if you want and have more control over the viewer.
there is churn in all major applications these days regardless of if it is needed. Reader is just trying to regain mind-share back from people viewing pdfs in browsers.
I also realized that there might be some pressure to turn into the next TurboTax, where the company eventually lobbies against improvements just so we can stay in business. I made a resolution that I'll never do anything like that. But I guess the founders of TurboTax never intended to do that either.
It was cumbersome on a Pentium 75MHz with 640x480x8 graphics. I'd be blown away 20 years later viewing those same files on a Macbook Pro with Retina display.
However, it was easily the best way to distribute printed documents EXACTLY the way they were meant to be seen.
They weren't meant to be edited, modified, data extracted from...And adobe went from a quick, minimal viwer to a bloated security nightmare by adding 'features.
Luckily, 3rd party and open-source projects came to the rescue.
I know there's that arxiv vanity thing, which is cool, but most of the time I get the "sorry we can't render this" error message.
It's still not something you can rely upon.
You still can have proprietary blocks of info inside the file.
The lack of open source tools to manipulate the format is a major hindrance IMHO.
It is also a very space wasting as well when people only do a bitmap dump into the file for scans. Forms are also an area that is not open source either.
It has so many hacks and kludges, it would be better if it was trashed and we start with postscript again.
Again I would be happy with HTML if the major browsers actually implemented full CSS print spec that exists to cover what PDF does with print layout.
> That’s an important feature, just like reflow. But does the former make PDF World’s most important file format?
It's not just that PDFs are displayed the same across devices, it's that it's displayed the same over time. I can open a PDF generated in 1995 and it will look identical today as it did then. The same can't be said about HTML or Word documents generated in 1995.
With very narrow exceptions, that are even narrower than PDF publishers think, epub is and should be preferred even for technical or academic writings. (Kindle's kf8 format is just a repackaged variant of epub.)
From a developer's point of view, when trying to enforce that submitted files are strictly in PDF/A format, from what I can tell there isn't much more you can do than dissect the file looking for umpteen disallowed features.
Is there an ISO-compliance validator to anyone's knowledge?
It was natural to view Postscript files on NeXT and UNIX machines, and Ghostscript was already thing. What could be better than just using the "native" language of the printer? I didn't realize that this was not a common view, or even possible for most personal computers at the time.
I was also misinformed for quite some time about the internal format of PDF, assuming it was just PS wrapped up in a container. In a sense, this is true, but there's a lot more (embedded fonts, transparency, forms are just a few that come to mind).
There's also a lot less, as PDF is not a full programming language like PS.
The failure of epub and other HTLM based formats in this use case, IMO, is that their focus on reflowing to support any display and device makes them inconsistent and therefore impractical for replacing PDF based content.
Sure it's horribly long, complex, comes with vulnerabilities and different consumers have different behaviour. But given the constraints of machines at the time it was created and the wide range of requirements and usages it's pretty damn good and has stood the test of time.
The library is mainly focused on text extraction but editing is on the road map.
The digital equivalent of physically printing and re-scanning the document.
It's the only way to reasonably guarantee someone will be able to open and read a resume in its intended form.
It really truly isn't. Anyone on any device can reliably open a PDF. That is not true for Word files.
Funny enough, this thread made me want to look at my old resumes. I have a .doc resume from 2009. You know what happens when I double-click it? Nothing. I don't have anything that can view it! Windows 10 doesn't come with a preview tool for doc files. Chrome/Firefox can't preview it either.
I tried to send a resume in LaTeX. They said they want a Word document..
[1] https://gitlab.freedesktop.org/poppler/poppler/issues/463
[2] https://gitlab.freedesktop.org/poppler/poppler/issues/683
To say in 2019 that a file format without mobile support is the most important is moronic.
But arguably executable file formats (.exe, .so, ...) are more important overall.
I prefer a long JSON file, which i can read from, or write down it to file format i want, to a hard-to-extract format such as PDF.
What's hard about it? It's an open standard with libraries in every language known to man.
If you use it as a dumb wrapper for scanned images, that's going to suck. But as a way to store and faithfully reproduce nicely typeset documents with images - ie, to make an archival copy - I don't think it can possibly be beat.
Extracting tables of numbers from PDFs is also a pain.