I asked a friend, and apparently some journals are still printed on paper but no one buys them on paper. So the practical reason for the use of PDFs seems to be that people can print them at home to avoid reading from the screen.
One actual reason for non-adoption of HTML, that I can see, is that there's lots of tooling available for collecting, organizing and annotating PDFs. But it doesn't look insurmountable, since files with metadata could just as well be HTML-based, and there are already multiple tools exactly for annotating HTML, only they position themselves for different audiences.
The thing is ePubs and such aren't all that easy to consume. They don't provide standardised predictable layout, so you can't control how it will be presented to the reader. HTML and ePub only provide layout hints to the device or reader app. Even the reader doesn't have much control how the document is displayed. In contrast PDF provides complete, precise control over presentation. If it looks bad or is misleading due to layout issues or truncated tables and such, all of which are possible and can be very significant in scientific or technical material, it's all on the publisher.
I read a ton of PDFs, but then I'm a gamer and most tabletop RPGs are published in PDFs. In that case it's because many of them contain a lot of tables, forms and illustrations. They're also trying to create an atmosphere and express a theme, so layout matters. For a lot of material layout makes such a big difference to how material is interpreted that it is content.
Meanwhile, I'm browsing the web all day every day and yet to see what is there to fail in a simple text-and-images, black on white layout that most academical and technical papers need. Complex formulas and legend positioning is solved with images for now—and I'd guess that with some industry effort MathML and SVG could replace them. Discrepancy in the support of HTML formatting features is solved with standardization of the current state-of-the-art for some ‘paper viewing software,’ just as PDF is standardized. And regular browsers could switch on the ‘paper reader’ mode for the format.
IMO it's inevitable that fixed sizes for layout will go away, for the simple reason that it's not how screens work. The sooner it happens in tech writing, the happier the rest of my life will be.
Also I just remembered that one studio specifically makes online ‘books’ to control the layout on the screen: https://bureau.rocks/books/ (though the use of the scroll-activation is questionable to me). Of course, they likely had to expend considerably more effort than is needed for either regular pages or PDFs. Alas, their explanation of the reasoning is not in English—basically, they don't like how formats like Epub are displayed but don't even consider PDF: https://bureau.ru/books/manifesto/
Ok, failure modes. HTML does not have the concept of a deterministic page size and length or predictable pagination. In academic papers references, figures, footnotes and titles need to be presented in a deterministic way so that when text refers to a figure, or a footnote is referenced, etc, they exist in a predictable and know relationship to each other. If the text refers to a figure that is supposed to be on the same page, the author wants to be absolutely sure it appears on that page. HTML can't even guarantee that a given image will be displayed at all. have you never been to a web page where some images were not shown or just appeared as placeholder boxes? Device and image library and support limitations might lead to an image not appearing at all, or at an unpredictable relative size so it's not easily readable, appearing in a form such that details are not easily discernible such as with bad gamma or transparency settings, or with colours not shown with fidelity. This can actually change the meaning or interpretation of a result.
Information that belongs together, that the author intends and needs to be together, had better be together. None of this can be guaranteed by HTML or ebook formats. These formats can't guarantee what fonts information will be represented in, they can't even guarantee whether text will he shown in italics, bold, what relative font sizes they will have, etc. Reference formats are very precisely specified, but these formats simply don't support that level of precision.
So ok, you can say you don't care about these levels of precision, but your asking why other people do and these are the reasons.
Solved with references to the structure, e.g. ‘Figure 1.1‘ or ‘Section 2.4.‘
> If the text refers to a figure that is supposed to be on the same page, the author wants to be absolutely sure it appears on that page.
Why? Place the figure right after the text, or before it. “Supposed to be on the same page” is an example of thinking in terms of layout inherited from printed media. Delivering information doesn't depend on there being a concept of pages.
> have you never been to a web page where some images were not shown or just appeared as placeholder boxes?
Solved with a format that ships the page and images together, like Epub does. It's a basic requirement for sane personal storage-for-reference anyway.
> Device and image library and support limitations might lead to an image not appearing at all
Image support is solved with standardization of the format, as are many other formatting support features. If one wants to read docs on something less capable, like low-end devices, they take a conscious trade-off for the risk of something being unsupported.
> or at an unpredictable relative size so it's not easily readable
Don't quite see what kind of issues you mean here. Size relatively to the text is rather easily specified, and afaik normally honored unless the image is of some wild proportions like extra-wide.
But actually, I really wish that browsers and readers could zoom in/out on any image. I already do this frequently with Mac's ‘zoom’ tool. It's probably not even too hard to implement, I should look into finding or making a browser extension for this.
> such as with bad gamma or transparency settings, or with colours not shown with fidelity. This can actually change the meaning or interpretation of a result.
Interesting point, however I'm hard-pressed to see practical instances of this. When academic PDFs are explicitly made to be printed, apparently usually on office or home printers, I doubt that high quality of the result is to be expected. In this situation, authors should probably opt for rugged but dependable looks in most cases, with simple gray-color graphics and easily distinguished shapes, since even color print is not of much certainty. If precise color graphics are necessary then the paper is not fully readable in the printed form favored by many people, and I guess they will have to refer to the images on their computer—though even then I doubt many of them calibrate their monitors. So, I don't see much difference with HTML being read on a phone or a tablet an ebook reader, especially since phones and tablets have pretty good screens and support for HTML features. Finally, for strict reference, precise graphics can be delivered as attachments in a stricter format—just like it's done nowadays for large color visualizations, afaik.
> Information that belongs together, that the author intends and needs to be together, had better be together.
Again, put images after text that pertains to them.
> These formats can't guarantee what fonts information will be represented in
Why is that necessary? For math symbols, Unicode has several variations like Fraktur.
> they can't even guarantee whether text will he shown in italics, bold
Sure they can.
> what relative font sizes they will have
Subtext and supertext are respectively smaller and larger than the normal type. What else is needed? Do people write “let polar coordinates be denoted with type of 36 points”?