The incredibly complex layout of ‘a wall of text and occasional pictures’ is there for a reason. That's what the authors wanted, and only PDF is up to the task of representing such delicate formatting.
Besides I don't think this is meant to be a total replacement that everyone is going to magically use and abandon PDFs, it might just be a convenience within a workflow to skim and check the full PDF if necessary.
There's only one rule of sarcasm on the internet, if you want to be sarcastic, always end your comment with /s, otherwise avoid it.
To get back to the topic, I detected the OP's post as sarcasm, but I had to read it most of the way through it before I determined it was probably sarcasm, but I wasn't 100% sure.
Moreover, the claim that sarcasm can't be detected on the internet is extremely dubious. It's really not that hard. The person puts extra effort into showing how the thing that they are pretending to believe is an absurd thing to believe. This is not culture or age dependent.
"It's awful how those people are doing all that good stuff" is a statement that requires you to know zero context. An alien can detect that this sentence is absurd. Moreover, if someone actually believes something absurd, and presents it this way, it doesn't matter if it's sarcasm or not, since they aren't convincing anyone.
If I wanted to hunt for a hidden meaning in written text, I'd rather read poetry.
Correction: It's really not that hard for you. It IS hard for me, especially in recent years.
What PDF-s generated by LaTeX have, that so far web failed to replicate, is beautiful fonts for maths, consistency, legibility and overall beauty of the end result. They also have the advantage that they are self-contained, so if I want to archive one or move to another place, I just grab the single file and I'm done. With web a single document is distributed across many files and you need make a weird wget dance (which is easy to misuse)/install some obscure browser extensions/etc - not exactly user-friendly. Also, if you want to print the document to get away from the distraction of the computer, you're guaranteed to get at least as good a result as what you see on the screen.
Don't get me wrong. This attempt is better than any that I've seen so far. Kudos to the author for that. It's no small feat. But it doesn't prove that PDF-s are obsolete and that people favouring them are irrational.
The real problem exists because most people don't use a correctly formatted/structured PDF to begin with. I don't wanna think about all the problems MS word might cause here and probably violates spec-wise.
Vectorized PDFs also don't use embedded fonts and content flow instructions but a bunch of randomly sequenzed glyphs that make no sense without a very good OCR. Vector-based PDFs are usually garbage for automated usage, and they are not useful anyhow for assistive purposes (e.g. a screenreader or a converter that uses the DAISY format or similar).
So yeah, I'd argue that PDF is the wrong serialization format. There are standardized alternatives that would be easier to parse, communicate, and license.
And countering your argument about portability: MHTML and WARC formats are very portable, the former is the default format for the page save functionality of all mobile smartphones. They are a single file, containing all necessary resources to display the page.
Have you seen arXiv Vanity? - https://www.arxiv-vanity.com
I wish somebody could make an extension or repository system to store all these, and prompt you sometimes when on the sites.
By the way, have you checked out Migaku's vacation add-on? I suggested it to you a while ago.
During the initial pandemic period, there were news articles I wanted to share with useful information. But the pages were full of Javascript based crap that guiding them on how to find the information became a task of its own. RemoveJS was very helpful in making those sites accessible to them.
I used in the past to locally correct for broken links within intranet/CI pages. E.g. something correctably wrong with a gerrit link, or whatever.
The primary aim is to serve the community with the outputs we have, while we improve the coverage and fidelity of our generator.
And yes, using only the official sources arXiv has released for reuse: https://arxiv.org/help/bulk_data_s3
This is indeed a major difference with -vanity
Could you explain what you mean here? Who is "we" - I assume ar5iv? "while we improve the coverage and fidelity of our generator" <-- does this mean this is a temporary situation, and in the future multiple versions of the paper will be available?
There's actually multiple "we", since there are two institutions involved, and one foil character - I'm the only one responsible for ar5iv "the website", in a personal capacity.
The fidelity of the generator has the "we" of the team behind LaTeXML, the TeX-to-HTML conversion tool. That is in many ways the most important project to remember here, as that is what we want to actively improve to a point where it is "good enough" in creating HTML over the entirety of arXiv.
The institution hosting the website, and wanting to "serve a community" is KWARC, a research group at the university of FAU-Erlangen in Germany. There are all kinds of projects and services brewing on that end, which have interplay with the HTML data behind ar5iv, but are not directly on the site.
And as to all of us reading HN, I think we are actually interested in arXiv itself being maximally useful. And so is the ar5iv site - it's a temporary deployment, that really is aiming to reintegrate back into the arxiv.org site, and general infrastructure.
If/when that happens is unclear, but in the meantime there is a lot of improvements that can be made, both in what HTML can be generated, deciding what the markup of scientific documents ought to be in the first place, as well as gaining some insights for what new problems arXiv would encounter if they served HTML.
I can technically implement that, but I really don't want to, as I see it as crossing a certain line. Seeing ar5iv as a limited, constrained, service is a good thing - I think it clearly communicates that I do not want to compete with arXiv.
https://arxiv.org/pdf/2001.00888.pdf (19 pages)
https://ar5iv.org/html/2001.00888 (missing content)
This is no doubt a hard problem ...
Indeed - hard problem and a messy solution. I have no easy answers.
"We are usually at least a month behind the live arXiv article list.
Also, we can only serve papers submitted with their LaTeX sources."
Moderate excitement is probably warranted, though.
(maybe for as long as since 2004 -- but ironically archive.org could not access arXiv for some years https://web.archive.org/web/20041101000000*/https://arxiv.or... )
Most other preprint services, such as those based on eprints (and in other disciplines) have always accepted PDF.
Rather, PDF files are accepted for the (fairly small minority) of papers that are written using alternative editors like MS Word.
What other sites have these kind of URL hacks?
example:
https://www.reddit.com/r/gifs/ --> https://www.redditp.com/r/gifs/
(there are a few other similar reddit --> image gallery url hacks, just google for them)
I wonder how well it works with more complicated mathematical formulas containing greek or arabic letters (although the example in that paper look fine to me), or other non-ASCII scripts.
other than that, this looks pretty amazing
I would prefer if your site requires a paid subscription so you can incentivize people to annotate the content. For a paper author, making the paper terse and complex is more impressive but for the rest of us it is very tedious to decipher.
(Though, from ‘tell HN’, the poster is probably not the site author, right?)
What I was wondering - and still am - shouldn't it be possible to get to a "Pleasant" justified layout on the web in general?
The jagged left-aligned paragraphs are some of the first bits people point to when invoking "my PDF looks better". I definitely am not saying I did it perfectly, but shouldn't it be possible to get a good justified scientific article on the web? Why not?
“Looks better” only works in regard to justification when one is admiring a page overall. However, that's not how people actually read text. They look at words in lines, and at that time uneven spacing keeps tripping the eye up. I know this argument, but there's no way around this discussion, and that's all there is to say about it (so far). Nicely looking bricks of paragraphs won't make the eye glide smoothly over the holes.
Page layout programs spend some CPU time on fiddling the hyphenation until the spacing is even. Browsers can't afford to do that, afaik (not sure about currently, but that was the situation a while back). Moreover, from what I vaguely heard, the HTML specification defines paragraphs in such a way that browsers don't even have the freedom to fiddle the paragraph height—something about the height being the minimum for the text on hand, or something like that.
Even if you insist to keep justification on desktop (though I personally can see the holes clearly)—for the love of good, please disable it on phones. It's just a mess there.
I think I can see that, but it's almost there, which is why it feels like there has to be something I'm missing for it to justify "just right".
But yes, at the least you've convinced me we should have a separate theme that goes left-aligned, and possibly makes a number of other choices that maximize readability.
Since I'd still want the folks that want "as good as PDF", to feel justified for sticking around.
https://www.biorxiv.org/content/10.1101/2022.01.07.475366v1....
If you want to do URL munging instead of using the UI, just add `.full` to the end of the URL ;)
This doesn't require that the manuscript be submitted in any particular format either. There are humans involved in making sure the full text formatting from PDF is good.
Also, ar5iv may disappear very quickly, since I am unsure if it's more helpful or harmful. But I'll definitely lean on the public attention to keep asking arXiv to integrate an HTML preview for their articles. In the one-and-only arxiv.org itself.
Lastly, one difference that may ignite a curious debate is that ar5iv is committed to being MathML-native. Yes. MathML is the only markup used for math syntax, and you'll see it rendered directly, undisturbed, with Firefox today.
Over 500 million MathML elements in the full dataset too, pretty awe-inspiring.