https://filingdb.com/b/pdf-text-extraction
OTOH it's totally possible to make a self-contained HTML page without using a JS framework of the day. It's going to be way easier to consume than a PDF.
https://filingdb.com/b/pdf-text-extraction
OTOH it's totally possible to make a self-contained HTML page without using a JS framework of the day. It's going to be way easier to consume than a PDF.
I do realize how ugly PDFs are to work with (I wrote my own PDF/A generator for issue 2[2]). This is a Tagged PDF though, so you can extract text using standard tools.
To understand the mindset, have a read of the Gemini FAQ[0], specifically the answer to why not use a subset of HTML - and then read Issue 2[2] which is a hybrid Gemini+PDF polyglot, for people who don't like reading PDFs, which is apparently everyone on this thread :)
Issue 1[1] also moves beyond PDF, to try addressing some of the accessibility shortcomings by (a) prepending the content as plain text, and (b) recording myself reading the whole thing out and arranging the file as a polyglot MP3 and PDF file that can be played in an audio player as well as viewed in a PDF reader as well as a text editor.
A mini-FAQ to address some points elsewhere in the thread:
* No, it's not going to replace your blog or the web in general.
* Yes, it's an experimental art project / longitudinal CTF forensics tournament / weirdo personal blog.
* Yes, I'm serious anyway.
But I don't really know that your PDF website doesn't use some evil invisible PDF feature.
And I have to use a special Gemini browser to access Gemini pages. (Since an HTTPS bridge misses the point)
So why not use Dillo as my "Sane subset of HTML"? It is not hard to hand-write HTML that looks great in Lynx, Dillo, and Firefox.
Actually, it is. I love Dillo, but it's very limited. I like to make my images "fluid" using max-width and max-height attributes, and Dillo will not support those in any foreseeable future.
But again, I still love Dillo.
How do you create that demarcated space where PDF/A, PDF 2.0, and all other PDF versions can be mingled together, and there's no easy way to distinguish them?
Designers would thrive in a PDF environment instead of handing their designs over to implementation as it is now.
Maybe PDF is just the beginning and maybe a similar format can be thought up that addresses some of the concerns expressed here, and move over in time.
PDF is an open standard, which is freely available2, and stable. It has a
version number and many interoperable implementations including
free and open source readers and editors.
I think ease of copy-pasting is one of the coolest things about the document-centric roots of the web (along with the back button and hyperlinks; in other words, hypertext rules), although the modern web does break it (along with the back button and hyperlinks) in many places, so I can see where he is coming from. PDFs aren't the answer, though.I'm basically in agreement, but the author has a good point that PDF is obviously self-contained and self-contained HTML pages are not necessarily distinguishable from those that aren't. Perhaps we might have to revisit MHTML or embrace Web bundles as an alternative to PDF.
Note that PDFs can contain JS too.
That's why he says to use PDF/A, which can't contain JS.
Wait, why?!? When does it render? Who's supposed to have a js engine to do that? What version? How does it load dependencies? Is HTML and DOM carried along with it? So many questions.
Who? The PDF viewer.
When? Since about 2000 in PDF format version 1.3.
Dependencies? Hah, no such luck. You're stuck with ES5 and Adobe's crufty JS library. There is no HTML and DOM, there are however some pretty thorough PDF document bindings.
Basically in the PDF world, Acrobat Reader is Chrome and everything else is, like, Konqueror or something. Don't be fooled into thinking PDF is a small spec. It's not.
On the other hand, there's nothing stopping you from using a double-barrelled file extension for denoting this sort of thing, e.g. "memex-opus.pub.html"; so long as it ends with something recognizable, double-clicking should still open it in the browser across all the usual platforms, AFAIK.
(I'm fond of using "xyzzy.app.htm" myself to take advantage of this trick for distributing simple, self-contained programs that are designed run in the browser.)
Completely agree. For instance, NASA's APOD site[1] is a good example of something that'd be nontrivial using both an offline PDF and modern lightweight alternatives like Gemini, but works really well even without fancy modern design. Under 300kB including the image (HTML's under 6 kB) before gzipping.
The author is obviously making a statements, exploring ideas... not searching for an actual solution to his use case.
The actual quote was from JFK iirc regarding the Apollo missions...
> “But it’s just as easy to write self-contained HTML pages!”
> Sure, but if you’re going to hide CTF forensics challenges in your publication, a coverdisk allows you to do it in style!
I think it's not meant to be taken extremely seriously