Portable Web Documents – An Alternative to PDF Based on HTML5 (2019)
getpolarized.io
getpolarized.io
EPUB lacks annotation, which is not at all trivial; I don't know how well it handles precise layout and pagination, and long-term preservation (will I be able to read it in 20 or 50 or 100 years?). But, again, I wonder why they don't improve on EPUB or at least use it as a starting point, and not seem to reinvent the wheel? (Sometimes, there's a good reason.)
https://en.wikipedia.org/wiki/EPUB
At a quick glance I don't see support for media formats like videos and such - in which case the obvious on-disk format is to hope the filesystem can compress text files and to layout a relative links only data-hive.
So the limitation is the limits of zip.
Features I need are being able to quickly open multiple documents side-by-side; a feature I actively don’t want is maintaining some sort of “library” for me.
Calibre's e-book viewer should cover your use-case just fine, and in KDE is quite easy to set that viewer as the default application for epub files.
I don’t want pagination for reference documents, for one thing – I want fast seamless scrolling.
Not sure if that’s still the case, but on macOS, it used to be hard to open more than one ePub file at a time using Calibre too – also an essential feature for research.
Finally, Calibre viewer edits each opened ePub by inserting a “last reading location” metadata file by default! It’s possible to deactivate in the settings, but I need to remember to do it for every new installation. It’s an absolute no-go as a default for a document viewer.
I don't quite understand: Highlight the desired files in the file manager and press enter? Tile the windows that open? I am missing something (obviously) ...
Microsoft uses that, of course. Why is it senseless? It seems much easier to work with. The portability of one file, instead of directories, is great.
I also love to get .pdf manuals of all the things I buy. I do not want to crawl through someone's website with chat popups and other engagement techniques to find out how to use or fix what I already purchased. Some of them are 40 years old, but as readable as new ones.
I also have lots of things like financial documents I've received that are identical and readable even 20 or more years later.
Although it isn't the cool-thing-of-the-day, I think pdf has its place.
It might be like phone/laptop users vs server users. Most people with a phone or laptop want the new best hot updated thing immediately. The server folks want nothing to change, ever.
> I make PDF files all the time with my scanner
In a hypothetical world without PDFs, wouldn't images work just as well for this? It's not like scanners are creating the same typographic layout etc. as a publisher needs to do when they send documents _to_ a printer?
I have some old ones, they were root document plus a directory full of loose gifs and jpegs.
I tried .webarchive but nothing can read it but the original browser.
Images of pages would be more of the same. You'd probably end up with a bunch of loose gif/jpg/png/tiff filee, one per page, or some semi-supported multi-image file.
Representing the pages as big images not only takes more space to store/archive, but also increases IO-time for loading the documents.
Then you can much more easily theme or site gen your PKB.
Not a huge thing, but enough to break the page layout.
HTML relies just as much on your monitor's colour grading as PDF does.
(Your other arguments stand well enough.)
The render as I wish. ie are readable.
The reader can alter its format e.g. change font and the result still works. (agreed that it needs to be decent HTML - but plain HTML works it is some of the complexities that make this not work and I consider that a bad HTML page)
PDF is fixed format and you can't increase font size to make it more readable.
However for some of those use cases then not being able to alter PDFs is a benefit - e.g. invoices and bank statements but then you probably need more than a plain pdf which can be edited (e.g. change a 1 to a 9) but a pdf with some integrity checks.
Many times even things like page or column breaks are extremely intentional. Having something "beneath the fold" or flowing onto another page can drastically change the way someone interacts with the piece of media. No PDF isn't great (my particular beef is with how there's almost no sense of a sentence, block, paragraph etc so it makes it almost impossible to copy or parse for text), but keep in mind that HTML/CSS only just reached near-parity in features in the past 5-6 years.
I'm talking about documents which are purely valuable for their content, not for their branding, and where the reader's accessibility is more important than the creator's design.
Reading on an electronic device is a different problem than providing something to be printed on a fixed format unchangeable piece of paper.
Thus you have different formats for the two different requirements. The issue is people keep trying to mix them up.
epub I believe are essentially what we as programmers would come up with, xml, html, and style sheets
disclaimer: this is only partially informed speculation
I'm not quite sure what you are saying: PDFs of course provide "properly typeset digital documents" to consumers. Do you mean that consumers don't need or want those? I think they do - look how much effort is put into presentation in every format on every platform.
I think, more importantly, PDF also provides consumers with a high-quality document they know they can read every easily, anywhere - any platform, any time, etc. - with just a click. What other format comes close? Imagine not having it - do you have the app to open the document? will it look like the original? when someone refers to the diagram on page 54, is the pagination the same for your copy/platform/etc.?
Also there is annotation, long-term presrvation, authentication (signatures), etc. but I don't want to sound like a broken record.
Actually that's exactly what I feel like PDFs don't give me. When I come across a PDF report, I immediately know it'll be a massive pain to read on a phone, on my e-reader, etc. and I'll have to wait til I get on a desktop. And even once I do, I have little control over the reading experience (I often use reader mode for websites to strip out daft decisions by designers, or even just to normalise font sizes to something I'm comfortable with).
Page 54 from the original is also trivial to preserve with any reader's pagination, it's just a numeric separator
> Page 54 from the original is also trivial to preserve with any reader's pagination, it's just a numeric separator
Pagination changes in different readers. That's why PDFs work so hard to be consistent.
That's what I said, but this doesn't matter much, you can have the original page 54 span for 4 reader's pages if the reader has a small screen with each of those 4 pages having the same number 54 so you can maintain the reference to the unmarked diagram
* User is on p.54, turns the page, and now they are on ... ?
* User tells a friend: 'Look at the second paragraph of p.54'. Friend: 'I'm looking, I don't see it'. User: 'Oh, the other p.54' ...
But the same issues exist today with PDF since you have page count (more visible since it's usually permanently visible in the app) and page labels, which are often different and thus can confuse your friend (by the way, does 2nd paragraph count include the last line of the previous paragraph at the top or does it start count at first full paragraph?)
(The better solution of more precise, e.g., per-paragraph, marks is, I think, also easier outside of PDF since PDFs don't retain text structure, only its visual position)
I don't think most end-users will like that.
> But the same issues exist today with PDF since you have page count (more visible since it's usually permanently visible in the app) and page labels
Yes, not a good situation, so why duplicate it elsewhere? Anyway, if I say 'page 545', IME people understand it's the document's page number and not the PDF page count.
But you don't need to duplicate it, all I was saying is PDF as exists has no benefits here.
Also, speaking of columns, when you reflow page to fit the narrow screen, you can simply make the same page longer instead of splitying, then you can have your same single writer page number. Basically, it's a non-issue for format comparison purposes, you can make UI whatever you like (unless it's PDF where nothing can reflow)
> when you reflow page to fit the narrow screen, you can simply make the same page longer instead of splitying, then you can have your same single writer page number.
You could do that, but IME, formats/apps besides PDF don't preserve pagination. Consistency across platforms, applications, etc. is hard.
In other cases, the information may not be destroyed entirely but much harder for people to see. For instance, automatically reflowing the London tube map would make it look totally unfamiliar.
epub is already a very widely used option ?
If it renders the same then it is not fit for purpose that is reading on a screen.
There are reasons to use pdf if the document must be the same on all devices but you need more than a plain pdf as you could edit a pdf to change 1s to 9s.
No technology is a universal answer. You want to understand the job and then find the right tool.
I suspect Fbreader might be OK - its earlier versions were but I have stopped using it after it added things.
Plus the basic unzip the file and just look at the files with a web browser.
These are just ones I have on my machine now.
> How do you do annotation?
Same as PDF - if you want that, get an app that has that feature.
> And how do you know you can read the document decades into the future?
Same as PDF, except that HTML is a simpler format.
> Also, for longer texts, isn't it harder to read without the high-quality readability enhancements, such as typefaces, layouts, etc.?
HTML has all these things? (With CSS, I mean.)
> Same as PDF - if you want that, get an app that has that feature.
In PDFs, annotation is a well-established part of the specification. I know my annotations on whatever app I use today will be fully functional in any app, on any colleagues or my computers, decades into the future. I could send you an annotated PDF document now and you could read it, annotate it yourself, and send it back.
>> And how do you know you can read the document decades into the future?
> Same as PDF, except that HTML is a simpler format.
PDF's spec prioritizes long-term preservation (or whatever term they use), especially in PDF/A format. HTML does not.
>> Also, for longer texts, isn't it harder to read without the high-quality readability enhancements, such as typefaces, layouts, etc.?
> HTML has all these things? (With CSS, I mean.)
PDFs look much better to me than HTML. I don't think HTML capabilities are on the same level, but I don't know the technical underpinnings of layout well enough to specify why. I know precise layout in HTML can be a challenge.
Uh, okay.
HTML is a dead simple plain text format that has been around for 30+ years. PDF hasn't been a well-supported open standard for half that. And it's no less complex now than it was in 2008, so good luck implementing a PDF reader on your own. In contrast, a mediocre programmer can hack something together that would make the contents of a given file in archive-quality HTML accessible within an afternoon. (Hell, you don't even need a programming system for it—any monkey with some patience could do it by hand, using a text editor while working on a copy of the raw contents of the file.)
(And you're not getting the "interactive charts" mentioned in the article without implementing most of that.)
The difference is what matters - The information or the look of the document.
HTML lets you get the information and you can alter the look. PDF locks in the look but is more difficult to get the information.
It's a question of the user's needs in that situation; to dismiss one or the other format or needs is just embracing ignorance as an easy solution.
However, if EPUB could get the precise formatting (maybe they already do), annotation, and long-term preservation correct, they could be all - or many - things to many people.
I think annotations also are not part of ePub they need to be part of a specification that build on epub. epub is the original.
Long term preservation is there as long as you don't try to do complex formatting in CSS.
It zooms better than pdf as the text will rewrap to fit the size of your screen whilst pdf zoom will force you to scroll;l to read a line.
Are you familiar with what the "/A" part signifies in PDF/A? Or what the words "archive-quality HTML" implies about an analog?
> writing a HTML renderer is an incredibly complicated task.
You're moving the goalposts. We're talking about getting the information out. You don't need to implement a full-fledged browser engine that's at parity with { Gecko, Blink, WebKit } to accomplish that.
Also, you're missing out.
Documents are best when they're just documents and can't instantly infect your device just by opening them. The odds of this format being any less toxic than PDF is zero. It'd be pretty nice if we could have a document format that printed well, but didn't also allow for RCE attacks, violate your privacy, or require a locked down environment and ad blockers just to view safely.
It'd be much better to support a sane subset of features than to allow a document to do everything and anything leaving you to just hope that the sandbox holds up to protect the rest of your system against the inevitable vulnerabilities.
these things are features in a lot of contexts.
I’m trying to read documents on my screen 99% of the time, not print them, so I couldn’t care less about some hypothetical paper page size they’re tying me to.
Regarding videos: PDFs support these too, and more – including scripting and 3D models!
Outside of that, PDFs are "hard" to edit (unintentionally) so for a quote or invoice I know you're seeing what I sent you. I've never had a PDF "updated" by a manager on my side, or yours.
And yeah, ultimately, they're often printed if they're legal or accounting.
If we want a format that prints terribly, but is great onscreen, and contains links, we already have that in HTML. If we want permanence and exact printing, we already have that in PDF.
I don't think it would be impossible to e.g. extend a variant of HTML (and ePub already supports this to some extent, I believe) to have a more stable print output (e.g. fixed pagination and page number hints) for when it's required, while still allowing reflowing, proper text search.
That's actually another pain point: PDFs is really a vector graphics format, so any text search is by necessity some more or less horrible heuristic layered on that. I work with a certain large company's PDF specifications every day, and they literally can't be searched in most PDF viewers, since the spaces between words seem to be done in a way not accessible to the heuristics of the readers, so everything is one large string to them.
If you want to cite a figure X from chapter Y, then cite figure X from chapter Y.
These aren't impossible problems. A well formatted index or table of contents gets you most of the way there.
But what if we used some sort of immutable reference? We could add new reference locations after some fixed amount of content such that they're never too far away, and we could even print these references at the bottom of each section for ease of use!
How does HTML hinder this?
Relying on your business partners or colleagues not knowing how to do this seems risky: If you can’t trust them, I’d personally use PDF signatures or at least hashes of the entire file.
I can't remember a single non-fiction epub that was better than or even equal to the PDF.
But anything remotely technical with code snippets, diagrams, screenshots, tables, etc, PDF is going to be infinitely better.
Javascript is not part of any document, and PDF does not have the equivalent - Printed pages don't change.
Right now you can download a webpage on most web browsers. It saves the html file and an associated folder. It would be cool if the web browsers placed that file and associated folder into a zip folder with a special file extension, and then the browser could read it on a double click as if you opened the html file.
Yes, I know it's limited. But dead simple. The format could evolve and get more bells and whistles over time.
Others mentioned epub. I do wish more sites offered epub downloads alongside or instead of pdf. Maybe browsers can do this, and maybe it'd be better than the simpler solution I suggested, but there are pros and cons to weigh here.
This browser addon does that for example: https://addons.mozilla.org/en-US/firefox/addon/single-file/
The only downside I can see to what you propose is the file size overhead of base64 encoding the assets.
[1] https://gildas-lormeau.github.io/singlefile-updates/version-....
Fantastic extension by the way, thank you very much for your work.
I always like to mention that that Chromium currently allows saving to MTHML (a self-contained single HTML file, with attachments/images). In Vivaldi (Chromium based) the appropriate launch argument (`--save-page-as-mhtml`) is enabled by default.
Firefox used to support MHTML as well until 2017 via an addon. Now days on FF an alternative called SingleFile saves to a literally renamed zip, for a similar purpose.
It's also guaranteed to never go away because it uses the exact same way emails are encoded.
My kingdom for an open standard (HTML) that can also be print-perfect (PDF) with good document header/footer/footnote and page numbers
I've been waiting on browser vendors for what feels like ages for good print-pages (since at least 2005).
And +1 for PrinceXML (https://www.princexml.com/) but i need tools that work with my MIT and GPL released code.
Yes you can - just save them locally?
All DocuSign contracts I’ve ever interacted with were signed using the “click here to sign your name” flow, i.e. not making use of PDF signatures at all.
I’ve signed PDFs using a “qualified digital signature” myself a few times, but that was in the EU.
no, its really not. I have been using PDFs in a "real world business context" for many, many years. and out of thousands of PDFs sent and received, less than 1% (probably 0.1%) had a use for a digital signature. signatures are useful, but you seem to be inflating their importance for typical daily business use.
I switched to SinglePage since I prefer to just save pages as stand-alone files. It does waste some space, but I can move the saved files to wherever it makes sense in my hoard and use my filesystem for organizing files instead of HTML documents living in their own location.
I also use the Save Text to File add-on. Often only the text on a page is useful anyway. At least that compensates a bit for the space wasted when using SinglePage on other pages.
> There are some associated file formats like WARC and MHTML that attempt to solve this problem but only really get you about 30-50% of a complete solution.
The whole point is that they are not dynamic.
Can you send me that as a "Pee-Double-You-Dee"?