I already can create PDF from latex, what does this repo add to it?
I already can create PDF from latex, what does this repo add to it?
Without special hints (that are ignored by regular PDF readers or printers), PDF is essentially a vector graphics format, and any of these tasks amount to an exercise in OCR.
This is a somewhat little-known fact about PDFs, since many viewers do in fact implement many of these OCR-like heuristics to provide features such as text selection, search etc. that make it look a lot like a text-based format, but it really is a vector graphics format at heart. PDF/UA makes this a bit easier.
As an example, consider a multiple column layout, as is often used in scientific articles. PDF-creating software not concerned with accessibility might just intersperse all columns line by line (i.e. present text in presentational left-to-right, top-to-bottom fashion), but it could just as well achieve the same outcome by drawing column by column in semantic order. Beyond encouraging that, I believe PDF/UA also defines a bunch of (invisible) metadata tags that readers can use to figure out semantic structuring of a document.
Seriously, I know of one experiment that has O(thousands) of papers as latex source, which they publish as PDFs and which (more recently) they convert back to markdown with some ML-based image to text engine. All so that it can be fed to a vector database to make searching easier, of course.
PDFs are a terrible format for machine readable information (and thus for reproducible science), but they are the currency of the scientific community and as such will be the standard output for the foreseeable future.
was terrible format.
PDF 2.0 Tagged PDF has MathML support, Index, Quote, BibEntry, Title, lists, tables, .... you can derive HTML from PDF if it's fully tagged.
PDF also support displaying and print production better than anything.
I'd much rather have plain text, or markdown, or asciidoc than PDF or HTML. It works everywhere. And support for embedded mathematical typesetting is a solved problem.
I mean sure, you can do it, but now you are on the other end of function over form.
I feel like PDFs are, for tons of usecases, the best we've got. The alternative isn't markdown, it's usually .docx.
Most viewers on iOS, macOS, Linux and Android that I've tried are very tied to the idea of all ePubs being ebooks and insist on managing them in their own library, don't optimize latency to first page rendered (because they usually index and cache a bunch of stuff for a "newly opened book", assuming you'll be reading it over the course of days/weeks) etc.
For example, on macOS, the default app associated with ePub is Apple Books, but to be able to use technical papers in ePub format, I really need something like Preview.app (and that doesn't support them). (In that way, the situation is very similar to JPEG XL: macOS supports them in quick view, but only Safari can actually open them...)
Not sure how it is supposed to look like on other OSes, but this is barely a GUI application on macOS. (I have to launch it from the command line, UI and text rendering are extremely pixelated etc.)
It's definitely not up to par with any PDF viewer I've used so far, especially when it comes to quickly navigating between pages and chapters (crucial for non-linearly working with scientific papers etc).
EDIT: there's also this https://sioyek.info/ found thanks to https://github.com/NixOS/nixpkgs/blob/9d13dbb1ef2d0a6a6788d5...
The Android version is a bit more than that however, it has a simple file browser, and I settled on it because it handles large files better than whatever alternatives I went through at the time.
For most content PDFs are used today, I'd much prefer extending something like ePub "downwards" with stylesheets, rendering hints etc. – that can all just be thrown out by an accessibility- or readability-focused viewer.
Yet, the PDF ecosystem is so large at this point, and so many organizations are still tied to skeuomorphic ideas focused around printed pages of text, that I can't see it happen anytime soon. I just hope that at least academia figures something out.
ArXiv's recent HTML efforts [1] are a great step in that direction, but as I understand it, they still use PDF as their main ingestion format.
[1] https://blog.arxiv.org/2023/12/21/accessibility-update-arxiv...