PDF: Still unfit for human consumption, 20 years later
nngroup.com
nngroup.com
The major attempts to replace PDF have largely failed, though. DjVu is relatively limited in scope. Postscript (as a document display format) has never been well-supported on Windows and is increasingly poorly supported on Linux due to rarity. XPS is perhaps the most direct "PDF replacement" but is nearly equally complicated (being based on the MS Office OOXML formats, giving it a similar cursed heritage to PDF's basis in the Photoshop PSD format), and there was never really a compelling argument to switch to it.
What I don't get is the suggestion that PDF should be replaced by HTML. The purposes of the two formats are basically orthogonal and replacing one with the other is doomed to failure. The author's argument seems more akin to "print-layout documents should be replaced by hypertext," and perhaps this is true in some cases, but it's definitely a different matter and one that the author's arguments don't really support that well.
In my opinion, hopefully more humble than the author's, PDF's main downside is the remarkable unevenness of the quality of the creation and reading tools, considering its supposedly "reads everywhere" nature. The "reference implementation" is a commercial product and supports a huge list of features that are rarely or never supported by third-party commercial or open-source implementations. The Linux toolchain still widely used with PDF (e.g. Ghostscript) is decidedly outdated and hard to work with, but there's not a lot of momentum towards development of more modern tools. All of these issues are likely rooted in the basic fact that the PDF format is extremely complicated, and so thoroughly implementing it is a massive undertaking.
The author's complaints about performance in particular reflect the flexibility and complexity of the format. Web browsers have mostly switched over to using pdf.js to render PDFs, which is completely satisfactory for documents that consist of text or images (like scanned documents), but can be absolutely unusable when dealing with extremely vector-heavy PDFs like GIS exports.
Even printing PDFs can become rather frustrating as the complexity of the format means that parse-related printing issues are relatively common. Even Acrobat, for a long time, would munge certain characters when printing due to some sort of inconsistency with how different generators and readers implemented font embedding leading to Acrobat not being able to locate the embedded character font. This seemed most common with the letter "l" but maybe I'm imagining that... but also maybe it reflects some frightening detail of the format or implementation behavior.
One of the most common issues around PDF consistency comes down to file size... different PDF generators are prone to create representations of the same document that are significantly different sizes. Scanners are often an extreme example, some combination of not "knowing the tricks" for PDF optimization and a probably very low-performance compression implementation means that low-end network scanners often produce PDFs that are hilariously large. Opening them in Acrobat and using the "optimize file" tool can reduce file size by 90% without apparent visual impact... the whole fact that Acrobat has an "optimize" tool (and that Acrobat Distiller used to exist) speaks to the scale of this problem. Inspecting PDFs that are "optimized" by Acrobat can be an alarming experience, as well. You may remember that this played a strange role in Obama's birth certificate some years back, as Acrobat seems to normally split PDFs into all kinds of different layers and apply strange transformations to them when it "optimizes." It's hard to know how much of this is actually "best practice" versus just a result of Acrobat accumulating decades of eccentricities.
So the bottom line is... PDF is too complicated for its own good, but then so are a great deal of other formats in widespread usage, like modern webpages which require complex parsing of multiple formats to render, and a great deal of historic cruft brought along with them. I'm not sure that there's any sound technical argument that PDF or web pages are a "better format," it's all a matter of opinion over whether you prefer print-format documents or hypertext, and that's going to be very application-specific.
Funny enough, I think one of the reasons PDF became so popular, is because it was originally seen as a "difficult / impossible to modify file that can be downloaded as a file and read in a static way". The lack of editing tools in the most popular PDF reader for a long time (Acrobat Reader) was the reason it became such a widely used format. Especially compared to distributing a .doc or .docx where the user can easily accidentally change something.
Isn't "the purposes of the two formats are basically orthogonal" actually the entire point the article is making? Literally the first line of the summary:
> Research spanning 20 years proves PDFs are problematic for online reading. Yet they’re still prevalent and users continue to get lost in them.
From the second paragraph:
> The [PDF] format is intended and optimized for print. It’s inherently inaccessible, unpleasant to read, and cumbersome to navigate online.
The bolded statement in the second paragraph that's clearly meant to be the One Important Thing to Take Away:
> Do not use PDFs to present digital content that could and should otherwise be a web page.
Your comment here is eloquent, but the article's argument is not "print-layout documents should be replaced by hypertext," it's "print-layout documents are a poor fit for reading on screen-layout devices." When you conclude:
> It's all a matter of opinion over whether you prefer print-format documents or hypertext, and that's going to be very application-specific.
Aren't you essentially restating the article's thesis?
I don't want to read an article online that's a PDF for largely the same reason that I don't want to print the web version of the same article rather than a PDF. It's generally going to be clunky. The print page size and dimensions are not going to be my screen/window size and dimensions. I certainly don't want to read two- or three-column text on screen, which may require zooming in and out and scrolling back and forth on the same "print" page. And God help me if I'm trying to do that on my phone or iPad mini.
The article isn't saying "PDF is terrible and nobody should ever use it"; it's saying "PDFs were meant for specific applications and in nearly all circumstances, online reading is not it."
PDF format is perfectly capable of holding structured text and other machine-readable/accessibility data, alongside the print-ready representation. Ask the developers of document authoring tools (starting with MS Office) and the various PDF generation libraries why they don’t include all that data as standard.
One can argue the appropriateness of a print-derived format in a constantly fluid digital world, but it does what it was designed to do and does it pretty well, and it could provide a lot more if developers and users could be bothered to do it.
And yes, HTML is its own exercise in awfulness that is equally bad at everything. I’d rather set my feet on fire that propagate that horror further.
Honestly, what we really need is a 21st-century Donald Knuth. I only wish she’d hurry up.
However, it is possible to turn HTML5 into statement-of-record documents (Not PDF's),and make them immutable, encrypted and authenticated. A HTML5 document can have the features we need from PDF (immutability, encryption, authentication, pixel perfect print, etc.) while still allowing the resulting document to be interactive and responsive (work well mobile & web) in nature.
Effectively the best of both worlds.
What I don't get is the authors' assumption that replacing with HTML means replacing with HTML that correctly uses "color, contrast, document structure, tags, and much more", leaves users in "a familiar context", is not "excruciatingly slow to load both on desktop and mobile", correctly employs "chunking, using bullets, subheadlines, anchor links, and accordions", and "show[s] a standard navigation", as opposed to … all that stuff that's actually out there. I don't know about the authors, but, if you give me a choice between a typical web site's idea about the flashy, JavaScript-heavy, animated, ad-laden way that I want to consume information on one hand, and a PDF on the other hand, then I'll take the PDF every time.
That's how I want it. I'm tired of new formats and new frameworks and new tools releasing every single year that pretend they're better. The best thing about PDF is that it doesn't change. Yes, whatever, Adobe adds new features, but any PDF ebook I download has nothing to do with that, nor that PDF scientific paper I downloaded this morning. I don't want random authors to be able to dictate how I interact with a medium, because the vast majority of them are, frankly, idiots when it comes to this domain and they have no sense of good user interface design.
You can, and I do, hope for this, and it's often what you get, but it's not at all guaranteed by the format. For example, PDF allows JavaScript: https://www.adobe.com/content/dam/acom/en/devnet/acrobat/pdf... . I think most readers other than Acrobat probably don't support it (which, of course, Adobe spins as a feature of Acrobat!), and not too many PDFs require it—although I do find that fillable PDFs, including from hopefully capable authors like the IRS, can be very finicky in anything other than Acrobat—but I fear it's only a matter of time until it becomes as difficult to "browse the PDF without JavaScript" as it currently is to browse the web without it.
> I don't want random authors to be able to dictate how I interact with a medium
On this score you're out of luck even now—if the author, intentionally or unintentionally, did a bad job creating the PDF, you're out of luck. This is one area where I'll give HTML the win: it is, or can be, good at things like dynamic reflow, whereas PDF not only isn't but, I believe, effectively can't be.
Improving the tooling will make PDFs load faster and (possibly) be easier to navigate. (Unless they're JPEGs stitched together in a booklet, in which case they're pretty much hopeless from a navigation standpoint in any event.) It won't, however, address the core concern, which you touch on here:
> What I don't get is the suggestion that PDF should be replaced by HTML. The purposes of the two formats are basically orthogonal and replacing one with the other is doomed to failure. The author's argument seems more akin to "print-layout documents should be replaced by hypertext," and perhaps this is true in some cases, but it's definitely a different matter and one that the author's arguments don't really support that well.
I mostly agree with you:
I think the arguments support it well enough: PDFs are sized for print, laid out for print, and fundamentally do not flow. HTML, despite the best efforts of some, is still malleable enough you can have a single page which gracefully resizes itself to a range of screens, a minor technical miracle the march of UX progress still hasn't fully taken from us.
My response to the author is this:
PDF looks the same on the screen as it does on the page. That is its blessing. That is its curse. Some people absolutely demand that as a hard requirement, and will not brook anything which adapts to different environments. If they didn't have PDFs, they'd go back to making websites where all of the content is in a series of JPEG images scaled to look right on their screens. I've seen it happen. Therefore, replacing PDF with HTML is not socially viable. It doesn't solve the "soft" problem, which is a harder constraint than any technical problem.
The meta-UX point is that a fixed layout allows you to use spatial relationships to indicate how fields are related. This is such an intuitive thing even inexperienced or amateur designers tend to do it by default.
Reflow doesn't allow this, and with some forms, reflow can literally make the layout - and the content - incomprehensible. There are workarounds, but it's often impossible to create a dynamic design that has the same mix of information density and spatial hinting as a static layout.
I can't help but think that HTML's third downside is the remarkable unevenness of the its quality.
The second downside is the remarkable unevenness of the quality of the CSS that is used with it.
The primary downside is the remarkable dangerousness of much of the JavaScript that is found bound to the HTML, that if turned off means that you often see a message that this site requires you turn on our dangerous JavaScript. At the best you end up back with the second and third downsides moving up a level.
on edit: perhaps a little facetious, but given the problems with quality found with websites that probably most of us are aware of it seems a bit much to complain about the quality of PDF. Maybe this is just some silly whataboutism on my part though.
Are you aware that PDF format has also featured JavaScript support (and related vulnerabilities) for a long time now [1]?
Just to clarify, the issue with the compression algorithm used by the scanner, not with the PDF format itself.
There a few videos of David Kriesel explaining the bug for the curious: https://www.youtube.com/watch?v=c0O6UXrOZJo
As far as I know, Firefox is the only browser that uses pdf.js. Chrome uses PDFium, Safari uses the macOS system pdf libraries, and Edge probably does what Chrome does.
I guess the extension to use pdf.js with Chrome is pretty popular.
Really though I might not be using the best term, esp. with the definition of hypertext being one of those things that's a little historic now. I'm mostly just comparing between print formats and formats where layout is done by the viewer to reflect user preferences (which is kind of a dead concept with HTML anyway, but...)
Is that really so? PSDs are through and through binary, while PDFs smell more like PostScript with extras...
It also has some optimisations for it's specific use case, such as that individual pages are completely described independently, whereas in PostScript the code generating any page can affect the content of any succeeding page. This is why in PDF files you can easily re-order pages or efficiently jump directly to and render any page.
The real advantage of PDF is that the images which are used inside the document are bundled into the same file... Whereas HTML has historically required the image files to be loaded from elsewhere which made it not portable. That said, now with HTML, you can define images with base64 data, so it could in fact replace PDF.
PDFs are treated that way, but it isn't really true. Due to the complexity of the format, there are many PDFs that will display differently in different viewers.
I don't see a problem with giving the document creator the option to go with a fluid or fixed layout and make the software default to fixed layout.
Except pdf.js is not satisfactory. Every now and then I come across a PDF file where text is invisible, because Firefox uses a blank font instead of an external font.
If it is a limited subset, I am okay with it. With increasing ubiquity of mobile devices, reflowing PDF is hard. I rather like EPUB these days. Which consists of XHTML. And reader support is also fairly ubiquitous with many Open Source software supporting Epub.
http://www.ecma-international.org/publications/standards/Ecm...
It doesn't help that nobody in Adobe could say no to every new feature that might possibly help to sell another upgrade to Acrobat.
Altough, for most people a non wysiwyg editor is too cumbersome.
I have yet to find an Android scanner that doesn't make pdfs that weight less than 350K per page. I have tried MS Lens, CamScanner, and a few others I cannot recall at this time.
Has anyone had success in this arena?
As far as a I know the svg format has comparable capabilities for graphics, all that is missing is a "page model" for html which would have to be invented.
But another barrier is that browsers refuse to support SVG fonts. One supposed reason for this, the lack of hinting support in SVG fonts, is less relevant now with high DPI displays - macOS no longer does hinting at all I believe. The additional effort to support SVG fonts is really minimal [2], so it seems strange that it's intentionally omitted.
[1] https://github.com/styluslabs/ulib/blob/master/miniz_gzip.h
EPUB. It's already zipped HTML and supports fixed layout.
Or, perhaps more simply, maybe just HTML + paged.js.
> all that is missing is a "page model" for html which would have to be invented.
CSS paged media is already a thing. What else do you need?
EPUB3 uses an XML packaging document, and the XML serialization of HTML5, so it doesn't really use lots of XML as opposed to HTML.
The world has also gotten more international and a PDF designed for US Letter Size doesn't fit A4 paper used in many other countries.
PDFs is left over from the early 90s when print was still the main way we communicated. We didn't yet email each other, at least not the masses. We didn't have lots of different devices. Our screens were low-res so it was much easier to read paper than screens (some people might still find that true). Heck, when PDF came out in 93 most PCs still ran DOS.
Now-a-days though does nothing bet get in the way. Sure, some rare PDFs can be reflowed but basically PDF wasn't designed for that and it's certainly not used that way. We need a format that re-flows for all the various devices we might be reading something on. For the most part HTML seems to fit that bill. Maybe a version with better image/diagram embedding would be good but we arguably do not need something brand new from scratch.
TL:DR; the world changed. PDF is designed for the the world from 30yrs ago. May it rest in peace.
I have built a business on PDF. I develop graphics software, enabling my customers to create large charts (36" x 96" and bigger) in PDF format, which they can take to the print shop for printing on large-format plotters and printers.
The sharp crispness of PDF text and vector graphics allows unlimited zooming while never pixellating (except the photos, of course).
If you are familiar with the technical specifications of PDF (1,300 pages 2006 ed.), you will appreciate the sophistication and power of the internal structure of PDF.
As an exchange medium, PDF has made huge contributions to commerce, technology and culture.
If the article title were "PDF -- unfit for web presentation" the author might have a stronger case.
You're quoting the click-baity title, but failing to actually read past the article summary's first sentence.
The summary's very first sentence states "Research spanning 20 years proves PDFs are problematic for online reading." This sentence alone frames the problem, and explains the whole point of it.
I don’t think the author was disagreeing with that either! They were saying, rather, that this effect is due to a collective failure to use HTML properly. If all you want to do is reproduce the physical pages of an article onto a digital device, and gain no more functionality (like text search, hyperlinking, reformatting), the author agrees PDF is great for that. But if you want to exploit all the features digital devices and the web offer, PDF constantly gets in your way.
You mean they look better on paper than on screen? And what magic improves quality of graphics?
Seriously html is used mostly for delivering spam and porn. And pdf's excel for technical documents.
Oh come now. I know we all like to hate on the modern web but this just isn't true. Wikipedia, CNN, Amazon, etc all use HTML.
You're just spewing a tautology. I mean, a well designed thing is still very good? Come on.
> And you can save a pdf locally.
You can also save epub and even HTML docs locally. That doesn't add much to the discussion.
> Seriously html is used mostly for delivering spam and porn.
It sounds like you're trying to force a morality-based argument to compensate for your lack of meaningful, rational points to make in favour of PDFs.
> And pdf's excel for technical documents.
They really don't. PDFs show good results in documents intended to be printed on paper following a very specific format, or whose main purpose is to deliver high-resolution vector graphics content intended to be printed.
Once your usecase consists of consumption with a electronic device, which involves delivering reflowable content that reflects personalized settings such as reader-specific accessibility settings and device properties, PDF fails to be an adequate option.
PDF is not the ideal mechanism for making a webpage or generally browsable thing. It's great for creating portable documents that look and perform the same over time. You can go into an archive in the UK and if preserved, read a legal filing submitted in the 1600s, and understand what it says. Likewise, if preserved, our successors will be able to look at digital PDF/A US Federal court filings in the year 2400 and understand what it says.
We already have content that is essentially lost from the 15-40 years ago due to file format issues.
PDF would have a similar problem, but Adobe leveraged their previous work on other products so they basically already had the rendering engine for Windows and it gained traction there.
Keep in mind that both Postscript and PDF were principally designed by Adobe. Adobe designed both because they were intended for different purposes, and this stands today.
Far as I can tell, most of the time (all of the time?) PDFs just seem to throw out this information, and manually select and place glyphs.
A work colleague worked on a document signing solution for a client once. Legally, at the time (and I hope this has improved), when a person added their digital signature to a document, that meant that they signed that exact version (read: hash of all the bytes) of the document.
That meant that PDF was sort-of problematic for the use case that the customer required: Giving the customer an A4 version to keep for their documentation was important - but having an A4 version on screen made for terrible scaling UX on mobile and tablet devices.
The fact that PDF is more than just text+formatting in that manner was a real hindrance at that point in time (2017).
(I'd be happy to know if I got any of this wrong after only hearing about it second-hand. This was in Switzerland, if this affects which laws were relevant at the time.).
However, in reality I've often found it easier to write throwaway perl scripts to analyse/modify PS files rather than PDF. Writing a tool to target a set of PS files made by the same process usually isn't too hard unless there is something really unique going on - but it can be a problem generalising it to any random PS file.
PDF is more structured so it can be easier to make general purpose tools, but in my line of work we prefer version 1.4 since the later versions add bloat that isn't necessary for print. It's also usually easier to consider them append-only, it's trivial to add content but editing is a lot harder due to offsets and references.
From my understanding, PDF is largely based on Postscript. PDF is to Postscript as HTML5 is to XHTML or HTML4.
The exact opposite is the case in my experience. Unfortunately the actual substance is often in a PDF and all the web pages pointing to it are superficial, copy and pasted and/or clickbaity fluff.
They then go on about how in web sites the content can be better structured and navigated. Unless I'm misunderstanding the word in English, what has that to do with whether the content has substance?
> [...] This leads to overwhelmingly long and inane PDFs
You mean something akin to a book?
I mean I suppose you could do all that with embedded JS, in theory, but one of the nice things about PDF is it mostly works absolutely fine with scripts turned off.
That stuff is thorough as hell and you even get schema definitions for all of it.
I'd pick that over poorly explained or 'discoverable' alternatives.
> and boring to read.
Not only is that subjective, but how is that relevant?
That's entirely cultural crap, and has little to do with either format. Or do you think that this HN comments page would be better distributed in PDF form?
The word "Unfortunately" is there on purpose. I often have to sift trough PDFs were it doesn't make sense to have the information only there.
What I'm disagreeing with is that PDFs unlike web pages lack substance. In my opinion the substance is often in the PDFs not because of the format, but how the information is produced.
E.g. within a government or enterprise the content could come from anywhere within the org structure, often multiple intermediaries away from the people putting stuff on the website. Everyone knows basic MS Word. On-boarding potentially hundreds or thousands of employees to a CMS and send them to a "how to craft effective digital content" which is what the Nielson article is ultimately selling is not always feasible. Only select pieces get a web treatment the rest gets summarized if not just linked. News papers have pipelines from Word to digital publishing tools to print / online because of this. But also this setup is not easy.
I just recently needed some specific information about traveling to and quarantine in Switzerland. The news sites where useless, the linked government web page was useless. Only the original PDF at the end of the link chain contained the information.
I'd prefer having this information more easily accessible/searchable. But as it stands, the substance is often in the PDFs, not the web pages.
Government reports, or those prepared by consultants, are often the worst offenders.
Didn't anyone notice that it's basically impossible to save an html page today and have it load and render correctly and offline tomorrow?
If I had an idea and wanted to communicate it, then I did so by recorded video, by live video, by blog post, by Twitter thread, and by HN comment, the same idea would be presented in very different ways.
In the same way, a writer who publishes something by HTML (blog post, etc.) will produce a very different document than if they intend to publish it by PDF (ebook, etc.). They tailor their message to the constraint and expectations of the medium.
We put all sorts of rubbish in the report to make sure it made a "thunk" sound when we handed it in. It was nearly 300 pages when it should have been 90.
The problem is, plenty will judge a report on its thickness. "It is thick, so it must be comprehensive." What percent of government reports are read cover to cover and what percentage are just ctrl+f through?
Perhaps it is better to be comprehensive in government reports than concise, to accommodate a variety of readers who want to drill into different aspects of the report.
(Of course, a PDF may not be the best structure for this! A well-formatted HTML reference with appropriate hyperlinks may be much more useful.)
http://media.metro.net/about_us/vision-2028/report_metro_vis...
I see it as the pendant of the "younger generations can't read anymore" critic, where lenghty, rambling and diluted prose is becoming harder and harder to parse and focus on.
On the other side page load speed and attention grabbing metrics are thoroughly studied for web pages and people value terseness, to the point of loathing click baits and endless listicles.
Even if verbose, those PDF are often still the only place where the relevant substance is together. The websites referring to them then cherry pick from it. I spend a whole lot of time sifting trough goverment PDFs over the last couple months because it was the only way to get to the information I needed.
It would be much easier if the content were available in different formats.
Screen size-adaptability and reflow remains a problem. It would be better to fix that on the PDF end than to move those uses over to inferior web technologies.
When you say "a multi-hundred page PDF loads in a blink of an eye compared to a advertising tracker-loaded web page", consider why that is. The basic reason is that every page in PDF can be rendered individually. (In fact, the top-level grouping in PDF is the physical page instead of the semantic model of HTML.) This is only possible because PDF has no layout! When you introduce client-side layout, the client must lay out every page to render any of them, because the locations of page breaks depend on characteristics of the client device, creating a sequential dependency. If you were to somehow add layout to PDF, the sequential dependency would be there too; there's nothing magical about PDF that would prevent it from inheriting the problems of HTML.
Finally, PDF does have animations and scripting (with multiple JavaScript engines). In fact, it even has 3D (old-school VRML-style 3D, not the flexible immediate-mode GPU APIs browsers have). You'd be amazed how bloated PDF is!
PDF is an atrociously bad format, and I don't know what "multi-hundred page PDF loads in the blink of an eye" for you but even a 100 blank page PDF takes nearly a second to fully load on my beefy rig (I did the test a few months back to prove a point). [Edit: Other commenters made the clarification below, but single page render time is not the same as document render time]
Clearly extracting text from a PDF is nearly as difficult as extracting it from a photo. Digitally extracting information from PDFs in general is awful, which makes the format awful for the various things it's used for.
Not to mention that many uninformed users today still install the garbage / malware PDF readers such as Acrobat because they don't know any better.
Sure. I agree we shouldn’t replace interactive web apps with PDF.
The manual for PGF/TikZ [1] is a huge PDF I frequently open. It's more than 1300 pages and has lots of graphics. It opens and navigates in the blink of an eye on my 3 year old laptop (with the Okular reader). PDFs aren't perfect, but they sure feel spiffy compared to modern webpages.
I do agree with some of the article's complaints, but not this one.
[1] http://mirrors.ctan.org/graphics/pgf/base/doc/pgfmanual.pdf
I will take the last part back, if someone can prove that I'm wrong about Word and PDF documents.
https://news.ycombinator.com/item?id=24035955
> No auto-play animations, no animations at all, no bizarre hijacking of scrolling
HackerNews commits neither of these sins. They aren't universal to the modern web, even if they're annoyingly prevalent. Given sufficient incompetence, both PDFs and websites can be bloated monstrosities.
I've never needed an ad blocker for a PDF. But I also don't have a good pdf reader for all of my devices.
Please remember that PDFs are absolutely capable of running code and do to deploy the advertising / tracking you listed as an issue with webpages.
If you are part of Adobe's premier advertising / tracking club (whatever it's called), and the user is viewing with Acrobat, you can see what people printed, where they highlighted, how long they stayed on a page, where they accessed, etc etc.
That's more of a problem with Adobe than PDF itself (never use Acrobat!), but that's hardly a rare theme when it comes to Adobe.
As annoying and obnoxious as animations and scrolljacking are, I prefer them to things like embedded viruses and SMB attacks that PDFs with embedded JS will happily run. https://www.sentinelone.com/blog/malicious-pdfs-revealing-te...
Normal PDF's are simple, reliable, and interoperable.
In contrast to webpages which are actually more often the "clunky", "slow", "stuffed with fluff", and "disorienting" (with scroll hijacking) alternative.
But the strawman is people creating PDF content as an alternative to HTML. Practically nobody is doing that. Virtually every PDF out there is designed to be a printable document first, that is then made available on the web. Nobody is saying "how should we architect our new site -- I know, let's make all our pages PDF's!"
What a truly bizarre article.
Tell that to Arxiv. Most papers never get printed. Everything is consumed on screen. Yet the layout is completely wrong for screens.
I think browsers should offer PDF reflow as HTML, to adapt to any screen width with optimal font size.
Web designers have the idea that I want a big column of text running down the center and lots of whitespace to the sides, perhaps with sub-menus. This would look OK if I had my main monitor oriented vertically, but I don't and almost nobody does. As a result only about 50% of my screen space is working and I am constantly scrolling back and forth on long pages if I want to look back more than a paragraph or two.
I've developed a deep dislike of commercial graphic designers as a class of people because they took everything that was annoying about magazines and put it on steroids. Many graphic designers hate text and now we have a million interfaces that look superficially interesting but are deeply unpleasant to read.
With academic articles, I virtually never want to simply read them online.
I need to save them for future reference, read them later when I've set aside time, annotate them, refer back to my annotations four months later...
Arxiv (or JSTOR or wherever else) is just where you get the papers. It's not where most academics are going to be consuming them.
(For consumption, a full-size tablet like an iPad, with a stylus or Apple Pencil, is absolutely ideal.)
Not at all. It would be bizarre if the uses of PDF that the article is meant to address didn't exist, but they do. Just look at https://berkshirehathaway.com for one example.
For a reference, we're a little over a week into the month so far. Yet when I check my browser history for PDFs, there are around 50 entries for August alone. Most of those instances are exactly what the author describes: cases where the format choice led to a worse experience than if that content had existed on a web page instead (or multiple ones). And as annoying as it is to try grappling with the format on a desktop screen, doing it on a smartphone would have been a non-starter, i.e. near 100% bounce rate.
My only objection is that I wish more of the content was in PDF, or at least had a PDF options.
https://www.sec.gov/Archives/edgar/data/1081316/000108131619...
When I download a 10-k, I dont want html to review on my phone. I want a PDF to read.
So ironically, this comment of yours describes your own original comment far more than the comment you're responding to.
Also, you do realize the Berkshire Hathaway site's PDF's are especially a lot of long printable documents, precisely what PDF's are designed for?
Finally, please don't assume bad faith ("relishing in spiting", "a crummy way to have a discussion") on HN. It's against the guidelines:
>Given PDFs poor usability for online reading, user-experience designers should either avoid using PDFs altogether in favor of presenting content on web pages, or, in cases where a printable PDF is needed, use an HTML gateway page.
But I strongly disagree that the article is arguing against a straw man.
Too many websites, especially from either very large organizations or very small ones, when asking "how should we architect our new side", answer "we already have some printable documents... let's make most of our pages PDFs!"
And TFA's point is: that's a disaster.
Restaurants create a "source of truth" menu that they print for use on-premises, often daily for seasonal fine dining. That is going to have layout.
Why is the restaurant going to go to the trouble of creating a second HTML version?
I've found that often, when a restaurant has an HTML menu on its site, it's months out of date because it never gets updated.
PDF menus always that are the same as what customers are getting at their tables, please.
The restaurant industry is poorly served by having these disciplines separated. It's been possible to do high quality printed output with HTML/CSS for a decade if you have a web designer familiar with it, but sadly too many aren't and so the restaurant has the menu done by a traditional print designer.
PDF also works GREAT as an archival format. I log into financial accounts regularly and save PDFs for each statement period. Makes reconciling a snap. And provides a locally archived document history for audits from taxing authorities etc. I never have to resort to finding paper.
Finally, PDF works great as a native format that my office printer/scanner understands how to write to. I can scan those annoying tax documents sent to my office to PDF and archive on the NAS/cloud backup as I deal with it and know that I have my documents digitized so I can shred the paper.
While this is just one use case of PDF framed in a browser, it still stands as one. I have also in my years regularly needed to archive the contents of a page - such as a receipt of a payment or a report on something.
In that case, printing the PDF seems to be one of the better practices. Saving as a Web archive (or whatever the format is called) is an alternative, but that is slightly harder to then print/fax or otherwise send to someone at a future date.
How is that a counter example? If instead of PDFs your bank had given you an HTML file encoding the same content, then it would satisfy the same purposes and have the other benefits that the linked article lays out.
I don't have a problem with PDFs myself, but surely it would be better if your bank gave you these in text form so that you can actually easily and reliably process them?
For eBooks, I've settled on reflowable EPUB. I guess, in some cases, we may want fixed format, where PDFs might be useful.
For online, I prefer HTML, usually as a continuous page, and with "pretty print" (@media print) CSS. I find it annoying that the page-break-% CSS rule seems to be ignored by just about every browser, or at least, interpreted badly.
I really have gotten a lot out of the NNG folks; in particular, Don Norman, but they do like to kick anthills.
In addition, typography and style are poorly adapted in EPUB format whereas with PDF - it can be read instantly on any device, any where and usually without installing a reader or fonts. There is so much inconsistency between Windows/Linux and Mac, iPhone, Android when reading an EPUB book.
EPUB = Great for text only literature.
PDF = Great for pretty much everything else.
I do a ton of academic research, and whenever I find a source in EPUB I have to convert it to PDF first, just so I can do highlighting, circling, etc.
True, EPUB applications generally allow you to highlight, but those highlights are stored in the application. They don't live in the file. You can't export them to import them in another reader program. With PDF, all my annotations stay in the PDF itself, and appear in all fully-featured PDF software.
Until these issues are fixed, I'll keep enjoying my beautiful LaTeX PDFs, thank you!
IMHO Adobe has been a terrible steward of PDF. I have no idea why. (Source: Used to write print production software in the 90s. Some of my team went to Adobe. One was bored so banged out a PostScript clone hooked up to the newer image library in a few weeks. They all said the PDF libraries were garbage, everyone was afraid to breathe on them, no one was motivated to do anything better.)
Re NNG: Agree. +1 Don Norman, whereas Nielsen and Tog haven't said anything interesting in ages.
Agree. EPUB is a fairly half-baked solution. Reflowable PDF might have been good, but, I guess there was never any support for it.
One needful use case was (is?) variable data printing, allowing mass customization, like direct mail. Pretty much the same technical progression that happened in user interfaces going from absolute coords and static layout to dynamic layout managers.
The fix was so easy. Just retain some of the source document's meta data, eg this group of glyphs are a "paragraph". The PDF object model was explicitly designed for exactly this kind of extension.
IIRC Some specialty vendors had some goofy work-arounds, like a post process tool for manually marking paragraphs.
But Adobe could never be bothered.
The Books app on iOS can do this.
I know Foliate does continuous scrolling, but it's only on *nix platforms.
But when you need the formatting, PDF is wonderful.
The problem with HTML is that it's a moving target.
The characters were overlapping, no matter which reader I use, what margin/line/character spacing I set, until I adjusted the font size to two points smaller, and everything was right.
It turns out that whoever made the epub, turned every word into a single html element, with a fixed position!
So, it was like PDF in a sense that every thing has a position on the page. But it was like epub in the sense that everything's size can be changed (albeit within the "div" element). And the default size doesn't event work for the book.
I cringed so fiercely that I nearly deleted the book.
%PDF-1.2
1 0 obj
<<
/Type /Catalog
/Pages 2 0 R
>>
endobj
2 0 obj
<<
/Type /Pages
/Kids [ 3 0 R ]
/Count 1
/MediaBox
[ 0 0 612 792 ]
>>
endobj
3 0 obj
<<
/Type /Page
/Parent 2 0 R
/Resources 4 0 R
/Contents 6 0 R
>>
endobj
4 0 obj
<<
/ProcSet[/PDF/Text]
/Font <<
/F1 5 0 R
>>
>>
endobj
5 0 obj
<<
/Type /Font
/Subtype /Type1
/BaseFont /Times-Roman
>>
endobj
6 0 obj
<<
/Length 51
>>
stream
BT
/F1 48 Tf
50 400 Td
(Hello World)Tj
ET
endstream
endobj
trailer
<<
/Root 1 0 R
>>But I get your point. PDF are not optimal for sure. And personal mileage may vary, I'm very sensitive to having proper text spacing and so on.
One of the things I absolutely despise about the rise of JS is that many many modern sites won't display anything (just the white BG) unless I allow 3rd party scripts. Is displaying simple blogposts and other textual information with the occasional image or video embed so hard that one needs to load often multiple MBs of JS from multiple external sites?
It's infuriating.
Give me a link... I might visit it in a browser.
I have a book, and I'd like to display video clips. AIUI, that requires that the viewer has JS.
This is not just for one journal; it's for the dozen or so journals that I look at regularly.
The mathematics looks terrible in HTML, and great in PDF.
Figures usually look terrible in HTML, and quite often when you click on the action to zoom them, you get a choice of just one zoom factor. Plus, the caption disappears so it's easy to get lost. With PDF, you can select your zoom factor and maintain context.
PDF has fixed page numbers, so you can refer to material in the paper easily.
The fixedness of PDF aids memory. I can look at a paper I've not consulted in 30 years, and know that something I want is (say) at the top of the right-hand column just past the figure showing such-and-such. With HTML, I basically get lost in a stream that changes if I zoom the text (often required to try to decode poorly formatted mathematical symbols) or even change the geometry of my viewing window.
I can highlight PDFs, and add comments to them. This is enormously valuable in research work.
(La)tex-generated PDF files can offer mathematical representations that are not just clear, but elegant, and in a form that matches historical convention. HTML representations vary from journal to journal (which is bad enough in and of itself) and almost never match what the reader expects from standard textbooks and classic papers.
I suppose HTML has the benefit that it can be set up to adjust to the viewing platform, so I can try to read a paper on my mobile phone. Not that doing so makes any sense at all.
For me, it's an easy decision.
And, when reading a Pdf, you can print it and get exactly what we see on the screen. So the tooling is very straightforward for the reader, just click print. With other Web content, it's the browser trying to fit things the way they should on a page and it generally looks horrible.
Positioning elements coded in marked language into a page is actually a tricky thing. Until we have the tooling to magically make any markdown content (with images) fit nicely in a page, pdf will prevail. Any hint on a tool that can take my markdown and print out beautiful pages, without having to tweak a dozen params, please show me.
This is almost impossible to do right. Even browsers can't produce a PDF that looks exactly the same as a web page.
If you want to programmatically produce a PDF from a web page, the best bet is to load up a full browser implementation just for that purpose as any other simpler solutions would certainly break the results pretty often.
PostScript was implemented in laser printers and printer drivers output PostScript language programs when you printed from an application like Notepad in Windows. High-end illustration and DTP programs output their own custom programs instead of being limited by the program output by the printer driver.
Over time it became obvious that the programming language features of PostScript were not being used very much. Printer drivers typically output a fixed header containing some function definitions then they use these functions over and over for drawing the content of the page. What if these function definitions could be built in? Then the programming language capabilities such as loops and conditionals could be left out and we would still be able to do everything we're doing with PostScript. In fact the resulting technology would be even more useful because rendering a page can be done without implementing a programming language interpreter. Thus PDF was born.
PDF made perfect sense in the early 90's when it was designed. Page Description Languages didn't need to be burdened with a programming language because no one was taking advantage of the language features. But then came the World Wide Web. PDF was the wrong tech for the Web, and PostScript would have been perfect. PostScript has all the capabilities of PDF, but it is also a programming language, which means you can dynamically alter how you render the page based on where you are rendering it. Alas, Adobe's direction was already set, PDF was going to be the future and PostScript is obsolete.
In summary, PostScript was invented at a time when nobody needed dynamic features, and PDF was invented for a static world but then the world suddenly changed and needed dynamic features.
Why do you use the past tense? PostsScript is pretty much alive and kicking, and the blue book is still one of the finest programming references in 2020.
I guess there could have been a use case for people typesetting their documents in Word or LaTeX and then "printing to web", but PDF took that role.
> However, Adobe desires to promote the use of the PostScript language for information interchange among diverse products and applications. Accordingly, Adobe gives permission to anyone to: ...
Source: https://www.adobe.com/jp/print/postscript/pdfs/PLRM.pdf#page...
For example, Agner Fog's instruction tables are something I look at from time to time, and hate browsing that PDF file for the information I need. Similarly, software manuals as PDFs are really annoying to use - and I've written them!
But for research that needs to be referenced through other research in a bibliography, having concrete reference points relative to the length/start of the content is actually much more reliable than having semantic links to headings or a URL. I'll frequently find deadlinks in bibliographies, or missing webpages, or webpages completely altered and unable to parse from an illegible URL. Versus a page number, which may be in exact or slightly wrong, but is a good starting point rather than a dead end.
While its annoying to go to sites using PDFs that should clearly be a webpage, its obvious that PDF is good at solving some class of problems for certain people. The scientific community for instance has been slowly moving towards formats that can generate both HTML + PDF, but for many reasons related to its legacy of print publication PDF is king.
To come in and just tell these people they're wrong is the height of obnoxious design hubris. Between that and the boastful self-accolades, delivered in 3rd person no less, its hard for me to take this seriously.
On the technical claims: while I agree that PDFs are not ideal for many uses on the web, especially for current attention-span-of-a-fly web usage, they are great for things where I am willing to dedicate more time for an in depth look at the subject. For those cases the complaints that authors list about PDFs (linear access to information, lack of advanced navigation options, optimized for print (i.e., look best on a large monitor)) are not limiting and in fact beneficial.
And some complaints (slow to load, stuffed with fluff, jarring user experience) are just as, if not more applicable to most of the web. My 2c -- work in R&D likely skews my preferences in the direction of paper as an ideal interface :)
Further, the arguments the article makes are gibberish:
"4. Stuffed with fluff. PDFs tend to lack real substance, compared to regular web pages. When you’re building out a web page, you can visibly see how long it’s getting and how far users will have to scroll to consume the content. Methods of structuring and formatting digital content such as chunking, using bullets, subheadlines, anchor links, and accordions help users efficiently skim and scan sections that may contain the answers they seek amid long-form copy. However, in PDFs, those techniques aren’t always used and content creators tend to favor quantity of content over quality and formatting. This leads to overwhelmingly long and inane PDFs."
"PDFs tend to lack real substance, compared to regular web pages." Really? Really? That's the argument Jakob Nielsen is going with? HTML is magically better?
"However, in PDFs, those techniques aren’t always used and content creators tend to favor quantity of content over quality and formatting." In HTML, those techniques aren't always used! They often aren't used. And HTML somehow enforces quality of content?
I read PDFs on screen everyday and I don't see the problem with it. It's honestly a great experience.
However, PDFs on the Paperwhite don't make for easy reading. I could and have converted papers to EPUB which is much easier for reading but less good for studying, and the purpose of these PDFs is studying. Yeah, I can grouse about PDFs but it's a tool which I use.
By comparison, I check EPUBs out from the library and they are surprisingly pleasant to read on the Paperwhite.
Yeah, the article is about the web and I'm answering about the Paperwhite. Maybe they have a point about browsing on the web. But for content meant to be read, for academic content, PDFs are pretty good.
BTW, on my MacBook I use Skim which is much better than Reader.
Not to mention on a tablet PDFs are much, much nicer to read.
And so alas it all sucks.
But I suspect HTML will eventually win this. While HTML can be printed, PDFs will always struggle with changing device sizes. Plus the web is becoming more of an app as time passes while PDFs will probably remain dumb content due to security reasons, so their applicable niche is growing smaller as the Web creeps in scope.
Simplified syntax, defined structure, not locked to physical form factor.
Still has warts, but a pretty good compromise.
I basically disagree with 80% of what this website says.
"PDFs tend to lack real substance, compared to regular web pages." made me chuckle. I don't know what kind of PDFs this person reads, but my copy of "Computer Networks: A Systems Approach" sure as hell has more substance and quality than a Twitter feed or whatever the author considers a "regular" web page.
One way to get the best of both worlds would be to have a normal webpage, but have the "Print this page" button generate a PDF that is nicely laid out. Often webpages are a mess to print.
I wonder how difficult it would be to write a tool that can turn a PDF into a usable webpage.
However in practice many websites don't put any effort into this, which is probably an indication that they wouldn't put any effort into a custom solution either.
> I wonder how difficult it would be to write a tool that can turn a PDF into a usable webpage.
I was looking into this but it is basically impossible. Since PDF is basically a collection of images (with some "fancy" stuff on top) you can get the basics, such as text and headings, however you won't be able to do much for semantics or layout. All web-based PDF viewers I have seen just render each page to an image and put invisible text on top for copy-paste support.
Look, web pages to the degree PDF is bad, is worse because it's riddled with adds and waits for servers you didn't intend to bother that,
- "help you" - "give you info you might be interested in viz adds - all other manner of social media nonsense - and to pay the domain owner $$$$
Maybe the OP should stop getting PDFs from NYPOST, Vanity Fair, Mad Magazine, Graphics Designers Guild dot com, or Madison 5th Ave S&Mrkting.com
Using PDFs to distribute content online instead of web pages is the real issue.
Same problem with trying to use a hammer with screws.
On the one hand, this DOM setup is crazy. Surely a more dynamic architecture would be better where decoders are downloaded on-demand as the user goes from site to site, coming across content types not yet seen. This of course raises security questions, as to the provenance of the decoder and less so of the content.
On the other, having come to the realization that the next logical step on the path to this scenario is WebASM, where the content and decoder are completely opaque to the user one can envision a world where there are a million different types of PDF, each with their own decoder, each trivially but crucially different. It's not a pretty thought.
ePub, mobi and such were developed to work around those limitation and make more usable book formats, but no web browser has native support for them (Edge had some support, not sure if that still there after the Chrome switch). Despite being HTML-based, those formats aren't really part of the WWW.
PDF does what PDF was designed to quite well, it's virtual paper. But the WWW has kind of failed to evolve into becoming a platform where you can publish long-form documents on, so PDF still continues to dominate.
I'm wondering what other file formats might work better, and why aren't they more popular? Epub maybe?
My understanding is that it's basically just a zip file with XML markup and any other assets like images. It's both human-readable and machine-readable, which is great for everything from version control to search to conversion between formats.
Isn't it true that every software project -- and indeed every project -- falls short of what people may want?
So for example, a word processor may be set to produce two-column text, and for paper that makes sense ergonomically. But it is horrible in combination with scrollbars. The same goes for margins at the top and bottom of pages.
A typical word processor allows you to easily switch text to one-column mode or adjust the page margins, so with just a few changes it could render your document in a more online-friendly way. So when you save as PDF, it would be neat if it could include both renderings into the same document.
In this hypothetical world, the PDF viewer would then decide whether to render it in faithful-to-paper mode or in online-friendly mode.
So, use web pages for presenting information best presented in a web page, and use PDFs for presenting information best presented in a PDF, and don't use a PDF when a web page would be better and don't use a web page when a PDF would be better. But that doesn't seem to be the point they're making for some reason.
Don't get me wrong, I find many aspects of PDFs hugely frustrating. But many websites are just horrendous and a complete misery to interact with. If it's more than a couple of thousand words I tend to start looking for a pdf version.
PDF supported embedded type, vector graphics, and many other features long before the web browser could. Honestly, the issue with pdf is how documents are created (often via fake printer drivers that often compile/translate whatever you are printing to some pretty gnarly postscript).
Consuming banking data as PDFs is a nightmare. The bank I was working with seemed to have spent _some_ money on its website (Regions Bank in the US, if anyone wants to know, but just so happens to provide .ofx exports for 19.95/month starting from the month you sign up, but not generated for previous months, although that's tangential). Meanwhile, my local bank that at first glance from its website seems like it's in the stone age provides a PDF statement that looks like it was made in the 80s (all monospaced font, no graphics), but they also provide a .csv export for transactions with seemingly no limits on date.
The latter bank approach signals to me that data is in the format it should be in. No more, no less. The former suggests The PDF and a pretty web UI is the de facto standard for communicating tabular data when it shouldn't be.
I get that PDFs online are a great alternative as a document that was originally meant to be printed and mailed, but it is a poor substitute for consumable banking data.
1. They are a flat format. Why is this good? When you text search for something it can be found, vs. in HTML where you can search only a single web page instead of a hypertext graph- I mean what would a complete search even mean in HTML?
2. They are also hierarchical. I can print a hierarchical schematic and navigate through it by clicking on sheet-blocks.
3. You can view 3d renderings in them. Someone can save their solidworks document as a .pdf, and I can open it and zoom and rotate the view in acrobat reader. There is certainly no standard way to do this in HTML.
4. There are no ads.
5. I can send somebody the complete thing as a single file. For a web-site I would have to send them a zip file that they then would have to extract- it's just not as nice somehow, though in theory it should be OK. This shows up in microcontroller documentation for example. Usually the chip TRM is a 1000 page pdf, but the software is a bunch of HTML files (a web-site really). It's inevitably easier to get the chip TRM than it is the software documentation.
Actually in this particular case there is more- the software documentation is generated as extracted comments from source code by doxygen and it is usually crap. Pdf documentation someone actually wrote, so it tends to be better.
When you get HTML documentation, there is often not an index.html file. If there isn't one, which document do you open first?
6. Every documentation as a web-site system has their own navigation method, whereas .pdfs have acrobat reader or whatever. Even on web documentation that has something like a go to next page or section button, it's hit or miss if it works well. For example, the placement of the next button will vary from page to page, so you can't easily just page through it.
The Word DOC format had the problem of becoming unreadable every few years until you managed to splash out and buy the latest and greatest Microsoft Word and its associated version of Windows.
At least the PDFs remain legible pretty much indefinitely.
Please check it out if you are interested.
I couldn’t disagree more with most of the assertions in this article.
Despite being "clunky", they render much more reliably than HTML on the decades timescale.
Chrome for Android doesn't support to load PDF so it automatically downloads and opens in another PDF app. Reading PDF in app is fine but I can't copy url from app because it's already downloaded. Normally I can just copy url from link but some site like Google and Twitter uses link jumper so I unable to copy url.
Yes it's same as other file types that can't be opened by browser, but other file types are rarely directly linked. PDF shouldn't be first-class citizen in web.
It's like somehow the PDF generation process randomizes the order in which it populates tables, such that selecting by a user later is generally impossible.
Maybe it needs to be interpreted / extracted from the PDF source itself, but average user graphical selection of a table is out the window.
https://www.sans.org/reading-room/whitepapers/malicious/pape...
https://digital-forensics.sans.org/media/analyzing-malicious...
Malicious PDFs are still around, even today
To get decent typography, one needs TeX, and TeX produces PDFs, not web pages.
For all the PDF hate, 99% of the time the rendering is better than most web pages, and it actually works properly. I don't have scaling issues with it on hi-dpi displays/etc.
I've also yet to see a browser do proper sgml/svg graphics scaling of high density (thing multiple hundreds of MB) maps/etc that are common in PDFs.
The hilarious thing is that these monstrosities are created by “security” people and work only in Acrobat on Windows with scripting enabled.
I am seriously wondering which universe this person is from.
Yeah! This year I decided to support Indie journalism and help the environment by not having the paper edition mailed to me. Big mistake. They'd literally rendered the print version as PDF, and reading that on an iPad was nearly impossible.
Perfect for information discovery. Rich annotations and hypermedia features (external links, document-internal links, TOC) in PDF fix pretty much all issues stemming from this. All searchable (if the PDF has been constructed properly). Permanent, static structure vs everchanging, confusing messes of websites. The web is NOT QUOTABLE and unusable without advanced full-text search. Barely an URI remains stable.
> 2. Jarring user experience. PDFs look completely different from typical web pages.
Typesetting on the web is a clusterfuck. Subpar microtype. Font rendering issues galore, tens of versions of popular fonts purchased at different points in time from different vendors with differently messed up CSS font configuration settings. Fonts are not embedded, but hyperlinked. I want a maximum fidelity reading experience for large portions of text and classic formats, because familiarity aids navigating a complex document. There is no need for fancy styles and whatnot.
> 3. Slow to load.
Renderers differ in quality and speed. PDFs render lightning fast at acceptable settings and if you wanna tune for maximum quality, you can do so at the expense of slower rendering. Besides, it took 2.401ms to load the web page these points are writteen on, excluding content blocked by ublock origin. This point is delusional. A 700 page beautifully typeset PDF opens and renders in <<1s on my 7 year old laptop, and my reader will prerender pages to speed up navigation even more.
> 4. Stuffed with fluff.
The entire paragraph is invalid because PDF has all those features.
> However, in PDFs, those techniques aren’t always used and content creators tend to favor quantity of content over quality and formatting.
The same goes for most web content put out today.
> 5. Cause disorientation. Because PDFs aren’t web pages, they don’t show a standard navigation like a website would.
Document structure is clearly presented in tree on the right side if the PDF is properly annotated/hyperlinked and the reader has a TOC view (productivity tooling should have this). Websites lack this discoverability almost always, and if they have it, it looks different and works differently everywhere, creating disorientation.
> 6. Unnavigable content masses.
This has nothing to do with PDF and everything to do with the reader in use. A semantic desktop would index all file content, allow cross-linking between files using file:// or other protocols, and generally expose all content to a local or internet search engine. Google search indexes PDFs just fine! (Again, a badly constructed PDF may not contain text at all or broken text, but that's a generation problem.)
> 7. Sized for paper, not screens.
This is correct, and an advantage, because the web and most other screen content lack the fidelity of typesetting systems like LaTeX, ConTeXt, InDesign etc which each incorporate decades of digital best practice, and several decades more of typesetting knowledge.
It is an disadvantage in special settings, like on mobile, but even then, PDF text can be reflowed with appropriate software.
> Users Strongly Dislike PDFs
It's my favourite format for archiving documents, knowledge, and even website printouts.
https://en.wikipedia.org/wiki/PDF#Logical_structure_and_acce...
But do any of you know how to use Pandoc (or some open source or command line tool) to convert a PDF to something easily readable on a Kindle?
However, PDFs are still better than everything else.
When I'm trying to learn something that is not short and simple, linear is good.
Far too often when someone tries to present a long and complex subject via HTML, they don't provide an easy way to go through the entire thing in an order that is pedagogically sound.
It doesn't have to be that way...but it usually is. I'm not sure why.
Instead, they provide each page with a sidebar that links to other pages, turning the whole collection into a directed graph of pages full of dead ends and regions that have no links to other regions.
You reach some page where the sidebar links to X, Y, and Z, which are all things that depend on what you learned on that page and you are now ready to learn. If you follow the X link, you may end up learning all about X but may never again see the links to Y and Z unless you remember that a dozen pages back you saw them and purposefully seek them out. It's very easy to not even realize that you missed a whole major subtopic.
In a linear format, such as an actual book, a PDF, an EPUB, or even a plain text file, the author or editor makes a decision on how X, Y, and Z should be ordered. Maybe they decide X, followed by Y, followed by Z. Maybe they decide X, then Y, then things that depend on both X and Y, then Z, then things the use X, Y, and Z.
Different authors might pick a different ordering, but they point is they have to choose something. Whatever they choose, you just keep turning the page and you'll hit it all.
For a big subject, maybe you don't want to hit it all. I've seen math books address this by having a list or diagram in the front giving you alternate orders to go through a subset of the book if you just want to learn just a subset of the subject.
In theory, HTML should be great for this, especially HTML with JavaScript. You could have a page that lets you select from different learning paths, and then the JavaScript would put "Next" and "Previous" buttons on each page that take you through all the pages on your selected learning path. You could still have the sidebar links, but if you follow one the JavaScript could add a "Return to Learning Path" button so you can always get back on track.
But until more HTML authors put in the effort to provide a linear path through the material that books/PDF/EPUB/text formats force their authors to provide, PDF and to a lesser extend EPUB will remain the best option for most people trying to learn a long and complex subject online.
(I give PDF the nod over EPUB because most EPUBs do not have mathematical notation that looks as good as it does in PDF. I don't know if this is a technical limitation of EPUB itself, or of the EPUB readers I've used, or of the authoring tools used to create the EPUBs, or simply the authors didn't know how to do it right).
A good example of HTML authors putting in the effort is "The Feynman Lectures on Physics" online edition [1]. That shows you can make a website that presents a long and complex technical subject that works as well as a book or PDF, yet adjusts well to a variety of different screen sizes.
Randal Munroe, "File Extensions"; xkcd.com, 1301, 2013-12-09.
There are of course some technical limitations to PDF that would prevent them from being mad "digital first", but even just changing page layout and adjusting margins and spacing and font for horizontal display (as most of our screens are) vs vertical layout as one would read a printed sheet of paper, would make huge differences.
I for one actually compensate for that in that I have a dedicated monitor that is vertically oriented in order to read PDF documents. Better yet if you can do it on a very high dpi screen. But even that is not ideal because although I actually like print formats, standards, and conventions (like margins, spacing, and structure), it's simply not relevant or applicable in digital until we get A4/Letter formatted tablets or desktop screens that emulate physical paper … albeit even that, inadequately. Nothing can really replace the advantages of paper, at least not until we get paper thin displays that have zero measurable response times on pen inputs … i.e., likely never.