ArXiv now offers papers in HTML format
blog.arxiv.org
blog.arxiv.org
https://browse.arxiv.org/html/2312.12451v1
It's cool that it has a dark mode. Didn't see a toggle but renders in the system mode.
Overall will make arXiv a lot more accessible on mobile.
While it looks perfectly fine on a phone. Two columns layout looks terrible on a smartphone, the text is too tiny to read comfortably.
It would probably be even better if you can flip it left and right like a ebook instead of scrolling to allocate the content faster. But current design is good enough IMO. (Compare to reading a pdf on cellphone)
arXiv accepts various flavors of TeX, or PDFs not produced by TeX [0], and automatically produces PDFs and HTML where possible (e.g. if TeX is submitted). In the case of the example paper under discussion, the authors submitted TeX with PDF figures [1], and the PDF version of the paper was produced by arXiv. The formatting was mainly set by using REVTeX, which is a set of macros for LaTeX intended for American Physical Society journals.
[0] https://info.arxiv.org/help/submit/index.html#formats-for-te... [1] https://arxiv.org/format/2312.12451
In the arxiv you use latex and do everything yourself. There is no editor.
Some fields may use Word files, but in most of physics you would get laughed at...
It is true that most journals will typically reformat your .tex in a different way than is displayed on the arXiv.
Two columns is good, albeit annoying on mobile. But the font. The typeface kills me, and almost every LaTeX-generated document sports it.
:root, [data-theme=light] {
/* --text-font-family: "freight-sans-pro";
}
it switches to "Noto Serif" that is way easier on the eyes.> "Computer Modern" is used for body text to give it a professional/academic look
I cannot find anything relevant in any of the 3 browsers I use (Vivialdi, Firefox, Chrome). Would really appreciate this option.
A quick search gave some apparently unmaintained browser extensions, and it's it.
The extra column next to the one I'm reading introduces a lot of visual noise, and the content is hard enough as it is. I'm sure physicists have all gotten used to it, but it certainly trips me up.
Papers are generally not read start to finish in one go: there's lots of rereading and jumping back and forth between key parts, and anything that moves them further apart makes this harder.
I still think a flexible layout is best. If you like multi-columns and have a wide screen, why not display 12 columns next to each other?
With PDF this is not possible. With HTML the content can in principle be sliced and diced how you like it.
But HTML is so much more flexible, and ideally people can choose how they want it, although at this point it seems that's not (yet) implemented.
I find jumping back and forth is always a pain on computer screens and ebooks by the way, and is the major reason I much prefer print for this type of thing.
(For reference: I am at the end of Gen X, people 3-4 years younger than me are considered Millennials).
I do not much care what font the auctor finds pleasant to read, but what I find pleasant to read, and this font isn't it, and neither are the colors.
Defaults and UX rule the world. It’s unfortunate that $subj wasn’t a thing for so long and probably scared millions of curious minds from material. It is so important.
The html version is wasting a lot of space on the right side and the color scheme is awful (dark grey on a brown background, seriously? How is that any better? Edit: disabling dark mode yields a better reading experience wrt color scheme). Also, somehow links to references make another http request and have no backlink?
The html version could make sense if it had more dynamic functionalities: change fonts/line spacing, toggle color schemes, maybe a mini map or some other navigational tool? Also, some kind of support for highlighting and/or annotating?
- I can imagine authors feeling frustrated if someone reaches out about a problem in the HTML version of their paper, but they have no way to correct it except by hoping that a change to the PDF fixes a change to the generated HTML. Easier to just fix the formatting problem in the PDF outright.
- It would be neat to allow people to experiment with alternative formatting for their papers. For example, imagine a paper about a programming language that embeds a sandbox you can use to play around with the language under discussion. Or a paper about multivariable calculus and you can interact with a three dimensional plot of some function.
What about various HTML tags that remote load resources? From script, link, to things like img or CSS `background-image` attribute, added in a `style` attribute.
There is a bunch of ways to do remote requests even without HTML.
But it is fine!There are no issues with arXiv generating the HTML and sending that over: they control the generation process, and users who visit arXiv already trust it to not be malicious. The issue is with letting the user upload their own and having it sent on to other users as is.
Please don't. Then you will have a mismatch between the source and the "own html" which ruins the point of uploading the source.
Enough!
My proof: https://scholar.google.com/citations?user=5DdrMc8AAAAJ&hl=en
I looked up the submission formats, and it looks like if you authored the paper in TeX/LaTeX, they do not accept pre-rendered versions of the document.
https://info.arxiv.org/help/submit/index.html#formats-for-te...
But if you did not author it in TeX/LaTeX (e.g., Word, Google Docs, etc.) it appears you can upload a PDF or HTML yourself.
With "sideloading" of HTML there is no way in general to make sure that the contents of LaTeX (and PDF) on one side and HTML on the other side is the same.
Is it not possible to write LaTeX code that produces different contents in HTML vs. PDF?
However, bugs get fixed, and since the PDF and HTML are generated dynamically, any such hack would be extremely fragile.
And while "single source of truth" can help to prevent such malicious discrepancy, it's unlikely that people would try to hack the system this way: what for?
Far more likely scenario is unintentional discrepancy, and single source of truth definitely helps to prevent that!
Yes, it is indeed possible to write LaTeX code that produces different contents when compiled to HTML versus PDF. This is typically done by using conditional commands within the LaTeX document that check for the output format being used. These conditional commands can then include or exclude specific content based on whether the document is being compiled to HTML or PDF.
In LaTeX, the ifpdf package is commonly used to check if the output is being compiled to a PDF. For generating HTML from LaTeX, tools like TeX4ht or LaTeX2HTML are used, and they often define their own specific commands or provide a way to detect the output format.
----- It gives simple code that uses:
The \ifpdf ... \else ... \fi command checks if the document is being compiled to PDF. If it is, the content between \ifpdf and \else is included. If not (which would be the case for HTML), the content between \else and \fi is included.
The content outside the \ifpdf ... \fi conditional will appear in both the PDF and HTML versions.
Can you recommend a system I can use to compile my latex, while also making sure the html is going to look good? I'd like some kinds of css style @media queries to switch between certain parts of the layout, while keeping a single latex file.
I think CSS is also backwards compatible.
It is the JavaScript birs that change
Unlikely. But if so, you can provide the packages yourself: https://info.arxiv.org/help/submit_tex.html#wegotem
> What if your paper isn't written with LaTeX?
Then they still accept PDF or HTML. See: https://info.arxiv.org/help/submit/index.html#formats-for-te...
(I'm still a fan of printing the PDFs, because I annotate on paper and refer to page numbers, but the HTML feature is in addition to PDF download, not a replacement.)
One thing that still sucks (not ArXiv related though) is reading mathematical formulae on the Kindle - wonder if someone with rendering expertise could have a look into the MOBI format.
Related projects:
https://github.com/ahrm/sioyek
https://github.com/arxiv-vanity/engrafo
https://github.com/dginev/ar5iv
https://academ.us/article/2111.15588/ (powered by https://github.com/jgm/pandoc I believe)
You can make decently accessible PDFs but it's lots of work, you need Acrobat on the producer' side and might also need it on the consumer's side. Free tools don't even come close. There's also the fact that the process of making accessible PDFs in Acrobat isn't itself accessible.
With that said, the way screen readers treat HTML math certainly isn't perfect, it's geared more towards school children than anything above calculus. I'm probably going to stay with my LaTeX source files for now. At least ArXiv offers those, not many sites do. To be fair, that approach also has its own set of problems (particularly when people use some extra fancy formatting in their math equations, making the markup hard to read), but I find this to be the best approach for me so far, at least on AI/ML papers.
Say I have Equation \ref{eq}. Why can't I just say "plot \ref{eq} for x from -6 to 11" and get my graph?
And yes, I know about pgfplots, PSTricks, TikZ etc. But in all those cases, I need to define the same equation twice, in different syntax to boot. It's kind of unsatisfying.
Pretty much for the same reason you cannot press a word and get a pop-up dictionary definition in a paper book.
An often cited example: what is f(x+y) ? Is it function f with x+y as its argument, or constant f multiplied by (x+y) ? TeX gives you no clue.
Or what is this i in your equation? Is it an index variable, or a square root from minus one?
You as a human figure this out by looking at the context and using domain knowledge. So does a "TeX to HTML/MathML converter". It is ultimately built on heuristics, and cannot be otherwise.
That's why I said basically "for the same reason a paper page is not interactive". It was designed this way!
The goal of TeX was to generate beautiful printed page. The need for semantic structure was not anticipated. To do semantics you need a "semantic version of MathML", or a language used by Wolfram's product, etc.
A simple example is ‘\sum’ which provides no way to capture the expression being summed over - because that’s not necessary for typesetting. That’s not the case in, say, MathML.
Writing MathML is no fun though because mathematical formulae are visually ambiguous and we rely on the context to know how to read them, e.g. does ‘f(x - 1)’ mean function f called with argument x - 1, or does it mean variable f multiplied by x - 1?
The amount it scrolled probably depended on the aspect ratio of the window, so it might be multiple key presses to scroll an entire column.
I would assume that the majority of persons on HN are not looking at their keyboard as they type.
I'm not deeply familiar with the state of that art, but it seems like recovering the metadata from a PDF generated by LaTeX would be no more impressive than many other things we're currently seeing language models achieve?
It was not designed to provide semantic information, unfortunately. So getting anything other than visual representation out of it is hard.
They argued that PDF was superior because the publisher could control how it looked and it looked the same everywhere but the point is that it should not. Things such as font size and line spacing should be at the control of the consumer, not the publisher. This isn't simply blind people but for instance also persons with dyslexia who use particular fonts to make it easier to read for them. Or in my case, someone who simply gets a headache from fronts and line-spacing that is too big. I've also been using darkmode everywhere for so long now that reading black text on a white surface on a screen gives me a headache.
https://tex.stackexchange.com/questions/485593/how-to-write-...
This problem is harder than you one would think naively.
The problem with this: you need to create a new standard, get everybody to agree to it, and get busy scientists who are concentrating on content and not representation to adapt this new standard in their writing, essentially requiring them to change their habits and spend extra time on writing (which many of them hate), for no obvious gain from their point of view.
I am not saying it's not possible, or not worth it, but it is not easy and simple either.
Besides, in HTML one can directly link to the relevant part.
Being able to link directly to the relevant part is irrelevant (pardon my pun!). Such links are machine-readable, not human-readable. Scientific text need visual citations and being able to name the referred part for reading comprehension.
And Harvard-style citations (AKA name-date) exist for a reason; when your read a paper even in interactive format it helps when you can recognize citations to certain papers and not having to click on them or memorize numbers.
Other styles have their own advantages and disadvantages; that's why they all exist and used by this or that journal, and no consensus on a single "right" style was ever reached.
[1] Great explanation here https://tex.stackexchange.com/questions/57717/relationship-b...
Anyway, if you (or anyone else reading this) has suggestions I'd really appreciate it!
This seems a massive gap in the market - many institutions have funding earmarked for such things.
What kind of turn-around time would be practical? Could you point me to any typeset mathematical braille that would be an example of a solution to your problem? Is Nemeth the only important standard, or are others important for you too?
I'm wondering if it's practical to set this up as back-office work here in Vietnam. There are some outlying provinces here where there are very few job opportunities. Job opportunities for the blind also round down to zero here (e.g. I could hire for proofreading). Maybe there's room to do something cool here.
Keep in mind that most blind people who speak English fluently but don't live in an English-speaking country (myself included) can't read English braille, or at least not well. Because of how voluminous Braille is, it uses contractions, single characters that replace common words and character combinations like "the", "would", "ing" or "ed". Those tend to be language specific, never taught outside their country or countries of use, and hard to get accessible electronic materials for. The math codes are completely different too, we use something derived from Marburg, while English-speaking countries use Nemeth. Even basic characters like + and - differ between those two, not to mention more complicated structures. It's not just the dot patterns that are different but also the design principles, like where you put spaces or when you can omit "begin fraction" / "end fraction" characters.
What would be very useful for me to be able to typeset myself are small things -- homework, quizzes, and (to a lesser extent) exams. Since homework and quizzes often have to adapt to what I actually covered in class, which may or may not match the syllabus, it's hard to rely on sending this out to be typset by others. (Exams are a little easier since they're usually done days ahead of the actual date.)
AFAIK Nemeth is the only standard that matters. If I can typeset a document, send it to the student, and they can get it on a braille display (no need for this to be on paper), it would solve a ton of problems.
Throughout college, my first question to most of my professors of math subjects was "do you do LaTeX, and can you give me your source code." Most said yes, and that's how we worked. LaTeX in, LaTeX or PDF out, depending on what the professor preferred.
The amount of LaTeX you need for calculus 1 isn't that great, you could probably teach it to a relatively bright student if you had an hour or two to spare, and then give them the source files. If you have the time, I'd suggest producing "stripped" versions of your files, with as little markup as possible to get your point across and no fancy formatting unless absolutely necessary. The amount of hoops some books and papers jump through to "look nice" drives me crazy.
You could also consider producing, teaching and consuming ASCII math, which seems like an even simpler and friendlier format. I couldn't really use it much in my school career for boring technical reasons, but it looks like a promising option.
One of my students was taking chemistry at the same time, which is (I think) much tougher for blind students. But they also had more teaching assistants for the course.
https://www.boia.org/blog/why-justified-or-centered-text-is-...
Having a two-column theme, or left-aligned vs justified themes, could be workable in the long run. I hope that we get to see some browser extensions modding the pages before too long.
The reason for the current justified text is that it is the default aesthetic for a LaTeX-based article, and a lot of authors expect it.
[0] https://info.arxiv.org/labs/showcase.html#arxiv-links-to-cod...
I really want journals to have two way links in a paper. I get google scholar alerts about certain papers being cited, and I want to skip to “why did they cite this? Did they use it, improve it, it just mention it?”
Thank you for the idea!
Unfortunately many institutions and businesses have ignored its limitation because PDF turned out to be an obvious-but-naive to put a 'sheets of paper' metaphor into a computer system, which in the 1990s appealed to tech illiterate folks doing bare-bones computerization of existing paper systems. So later we got complicated and error-prone tools for editing PDFs, and many random additions to the spec to allow for unusual use cases.
As an academic researcher, generally speaking I also prefer PDF, and the inflexibility and static nature is a feature, not a bug. I appreciate the fact that a paper will appear the same everywhere, that I can refer to "the top of page 7", etc.
The exception is if I wanted to just skim a paper; in this case, I think I'd prefer HTML.
I'm a huge fan of what arXiv is doing here. It effectively preserves the status quo, while adding an additional option on the side. The HTML option might prove a little bit useful for me, and it is likely to prove extremely useful for people with disabilities.
There are many great solutions to this problem, including ones that don't require Javascript at all. This website (https://gwern.net/silk-road) presents a really good example -- every header and sub-header is a clickable anchor. If more granularity is needed, on newer articles most of the paragraphs start with an italicized margin note -- though for technical writing, paragraph anchors might be better. The page also pays careful attention to print CSS and has a 'reader mode' to convert all links to footnotes when printed.
Some websites will also preserve the text you select in a URL anchor, but more often than not this is just cumbersome. It also has a greater risk of not surviving changes to the webpage.
It's actually hit ~88% of the market https://caniuse.com/mdn-html_elements_a_text_fragments but unfortunately, Firefox remains a holdout* and that's my browser, so I don't use it (although maybe I should just install https://addons.mozilla.org/en-US/firefox/addon/link-to-text-... and try it out - my existing method of making new anchors for annotation purposes is cumbersome).
* Firefox officially is positive on it but no sign of any movement on it: https://mozilla.github.io/standards-positions/#scroll-to-tex... https://github.com/mozilla/standards-positions/issues/194 https://bugzilla.mozilla.org/show_bug.cgi?id=1753933 https://wicg.github.io/scroll-to-text-fragment/
The format in which articles are submitted and stored in arXive is LaTeX. PDF is automatically generated from it.
Probably arXiv does some caching of PDFs so they don't have to be generated anew every time they are requested, but I don't know how this caching works.
The default PDF format puts the xref table at the end of the file, forcing a full download before rendering can take place. PDF-1.2 onwards supports linearized PDFs, and most PDF export tools have some way of enabling it (usually an option like "optimize for web").
I love PDF.
IIRC, ar5iv was created on his own initiative by Deynan Ginev
https://twitter.com/dginev/status/1736792316675825981
and it seems that he has worked tirelessly to fix nearly all of the edge cases during the collaboration.
This project creates huge value to humanity so Deynan is to be heartily thanked.
1. My name is Deyan (hi!)
2. ar5iv was the latest frontend incarnation, but our actual work on converting LaTeX to HTML goes back nearly 20 years behind the scenes.
3. I was an undergraduate student when I was introduced to the project back in 2007. It was started "in spirit" by 3 senior co-conspirators back then: Michael Kohlhase, Bruce Miller and Robert Miner. And I am by no means a solitary actor today, even if I may be the chief online presence of the people involved. Bruce is doing the bulk of the hard work on LaTeXML to this day.
I documented some of the history in an invited talk for CICM 2022, which you can find on youtube, or see the slides at:
https://prodg.org/talks/welcome_to_ar5iv
It's really great that the HTML has now reached "home base" in arXiv, and I hope their team gets a lot more of the positive attention going forward - today's achievement is entirely theirs!
I had come across latex2html, Dan Gildea's project, and found myself unpleasantly dissatisfied with how it worked. As I understand it, it's more a "half implementation of lots of packages" rather than what ar5iv seems to be, which is "enough of the core LaTeX engine producing HTML instead of DVI"? I'd love to know more about the nitty gritty of how the engine does its thing.
I'm curious: How has modern web tech (e.g. WebAssembly, Canvas, etc) helped or gotten in the way of getting good LaTeX rendering in the browser?
Which also allows us (and generally all contributors of latexml package support) to conveniently maintain various parallel data structures and metadata needed along the way.
Modern HTML is very often helpful to produce higher quality article renderings. Examples:
1. we recently started using flexbox for subfigures, allowing them to reflow.
2. we have started emitting ARIA accessibility annotations (there is now an "alt" key for \includegraphics)
3. MathML Core allowed us to have native web rendering for math expressions in every browser.
As to LaTeX rendering in the browser, there are various other projects out there you could look up with partial support. For latexml the WebAssembly route seems most realistic, as we are undergoing a rewrite in Rust. But there are quite a number of pieces to flesh out before we get there.
Btw given we are into quotation academic world, I wonder whether you may have mention Gartner Group to invent that technology curve. To be honest there is a variation I like more which deal with the chasm issue.
For now it only works for papers submitted this month. But it's great to have this feature, makes it so much easier to read on phones.
I've written my thesis in Markdown in the past because of this (best for humans) which can be easily transformed to HTML, Word, PDF and even LaTeX https://github.com/tompollard/phd_thesis_markdown
And I think that XML is the best format for machines.
article {
text-justify: Knuth-Plass;
}Here’s a discussion of hacks to achieve the algorithm’s results on web pages and an upcoming CSS feature as of 2020. https://mpetroff.net/2020/05/pre-calculated-line-breaks-for-...
At least the HTML version pairs each author with their affiliations, instead of the PDF which has all the names on page 1, and all the affiliations on page 2. That's completely unreadable.
It's highly unlikely anybody will read an entire author list this long; typically you would read the first two or three names, or check if some particular name is on the list. So the compactness of the list and being able to quickly get to the article contents is important.
> Default to HTML: HyperText Markup Language (HTML) is the standard for publishing documents designed to be displayed in a web browser. HTML provides numerous advantages (e.g., easier to make accessible, friendlier to assistive technology, more dynamic and responsive, easier to maintain). When developing information for the web, agencies should default to creating and publishing content in an HTML format in lieu of publishing content in other electronic document formats that are designed for printing or preserving and protecting the content and layout of the document (e.g., PDF and DOCX formats). An agency should develop online content in a non-HTML format only if necessitated by a specific user need.
https://www.whitehouse.gov/omb/management/ofcio/delivering-a...
I think the general problem is that the end-user doesn't control an html document, e.g., for annotation, as a local record, etc.
What do you think of the epub format?
Despite all our advances, we lack an editable, local, multimedia, platform (and form-factor) independent, self-contained file - essentially a word-processing file for the 21st century (and I mean it's almost a quarter-century overdue). epub has that potential as a format, and being based on web standards it has capability, a universe of supporting tools and technology, and easy adoption to different applications.
But I haven't heard anyone else express that particular interest, and as of a few years ago epub doesn't allow annotations and is not stable (i.e., I don't know that today's epub file will be readable in 20 or 50 years) - two essential requirements for a serious local content, imho.
And even if it meets those specifications, we need epub editors that are the equivalent of word processsors for non-technical users.
Seriously, name a single device that has PDF support that doesn't allow you to view HTML.
I think you're conflating "html" and "things stored on a server", because all of your objections apply to pdfs stored on a server. The ability to save and annotate pdfs is not an inherent feature of the file format, they exist because the format is such a PITA to interact with that specialized programs have to be written. HTML can be saved just as easily, and usually is (on archive.org).
It's an ISO standard with a very large ecosystem outside Adobe. Many users and businesses I know don't use Adobe at all.
As far as annotations, you can use the native <ruby>[1] tag, or strikethough, but if you mean "literally drawing on the text" then, yeah, you're looking for an image format at that point (which is fundamentally what PDF is), but we shouldn't default to storing text in image formats just because of one specific use case. (Also, as I said above, the only reason tools exist to easily do that in PDFs exist is because everyone insists on using a format that's hard to edit. )
Also, note that the context I was responding to was US legal documents, not something more presentation-heavy.
1. Saving as "Webpage, Single File" (.mhtml): Neither Firefox nor Chrome even showed up in the list of available apps to open it.
2. Saving as "Webpage, Complete": Opened in Chrome but images were broken. Also very difficult to open with the default file browser because it uses a flat folder view and the sidecar folder pollutes the file list.
I was hoping this would work, perhaps you will have different findings. I agree that HTML is the superior format in theory but usability in practice is often lacking. I'm resigned to using both depending on context.
Such as? What doesn’t have a browser but can render pdfs?
We could have such a format if browser and os vendors were interested in supporting such a use case. Unfortunately, they aren't.
On the browser side, supporting all-in-one html files can be as simple a reading a single multipart-encoded page. Heck, if they support automatically serializing all external resources as datauris when saving pages, then most browsers will be able to open them without any modification.
On the OS side, operating systems can treat html files as first class citizens; execute them in an offline sandbox (most operating systems have embedded webviews), then extract icon, title, description and other metadata to present to the user. An icon the consists of a blank page with a small browser icon in the corner doesn't tell me anything about what the page is about. This needs to change.
In short, html can be easily made nicer to deal with locally thanks to all the parts already being in place. The problem is that no one (tech giants, os vendors) are interested in doing this.
Ctrl/Meta/Cmd + S should do the trick, or "File > Save page", and you get a HTML file you can open in any browser. If there is images, they'll most likely be loaded remotely, or worst case not load at all. But the rest of the structure is there.
Most sites have images as a relative path which won't work with saved html and there is also CSS.
Interestingly this review paper seems to have their side by side figures intact (e.g. fig 2 fig 4). Maybe it's because he used a subfigure like environment (judging by the subcaptions)?
Getting subfigures emulated via flexbox is one of our more recent LaTeXML enhancements, and still has some ongoing work (working on it today actually). It can be a bit finicky to test - there are easily 20 different ways people can write LaTeX for subfigures in arXiv.
At my institution, all of the lowest quality drafts I read are made with latex. I think it's because the programs people use to write latex do not have spelling and grammar checking. Also, the people that prefer latex, are the same types of people that are more interested in technical things, than spelling and grammar.
you can run toggleColorScheme() twice in console to switch to light theme or dark theme.
The magic of inline images at a known DPI, of course you can provide images for different DPIs.
Reading maths/science noscript/basic (x)html documents on my 100 DPI monitor, on wikipedia. Not yet fully ready on arxiv.
Latin Modern is used by:
- Wikipedia. - Math.StackExchange. - Nearly all papers, including the ones hosted on arxiv in PDF format. - Nearly any math videos, slides/presentations, notes. - Almost everything, really.
Palatino just looks weird.
Also, I imagine that authors might do math formatting hacks that were only tested on Latin Modern, and might end up breaking on Palatino.
TL;DR:
Palatino :(
Latin Modern :)
Edit: aaaand they got Fastly https://news.ycombinator.com/item?id=38723373
However, ar5iv isn't a la carte like arxiv-vanity. They pretty much do last month's papers every month or so. Something like that.
You can think of both arxiv-vanity and ar5iv as the "alpha" experiments that lead into the official arXiv "beta" HTML announced today.
Once a few rounds of feedback and improvements are integrated, and the full collection of articles acquires HTML in the main arXiv site, ar5iv will be decommissioned.
The plan is to turn all existing ar5iv links into redirects to the official HTML, and free up the resources for maintaining it. I am not sure what are the plans for maintaining arxiv-vanity, but I suspect they may head down a similar path some time later.
Reminds of Burning Man when people kept telling me, "Never talk trash on the art at the main landmarks. The artists are frequently within listening distance."
So, of course, I'd walk around talking about buying the art for $50K-$60k, knowing it's already scheduled to be burned with the landmark.
The most versatile tool I know of for converting various document formats, including PDF to HTML, is the oss ebook tool Calibre: https://manual.calibre-ebook.com/conversion.html
I have seen https://pdfbox.apache.org/ used for extracting text from PDFs for analysis, but you won't get HTML output.
Also, a bug in a converter is conceptually much easier to fix than to re-train your LLM.
I am not sure that AI in it's current state is useful when "high fidelity" is required.
We have a plan in place to meaningfully fall back for unknown packages, but that will take at least another year to put in place, and likely another couple of years to stabilize.
Meanwhile, there is some hope that with arXiv launching the HTML Beta we will get more contributions for package support (LaTeXML is an open source project, with public domain licensing, everybody benefits).
But again the original point is spot on. Coverage will be hit-or-miss for a while longer yet, for an arbitrary arXiv submission. The good news is that authors could work towards better support for their articles, if they wanted to.
For example, HTML isn't divided into numbereres pages while PDFs are. A lot of latex interacts with page boundaries. Figures tend towards the tops of pages. And there's \clearpage. And the reference list might say which page each citation appeared on. All that stuff needs someone to decide how to handle it and then to implement that handling. Like... what value does \pageheight return? Sometimes I resize things to fit the page height, and if it was doubled then I should have resized to fit the width instead.
It's nontrivial to export this to HTML in all cases, and even then, nobody is asking for HTML from us even though we all want it. I'm guessing Arxiv is using some kind of converter which _usually_ but not _always_ works.
That said, this is a long time coming and PDF as the standard should've died a decade ago. I wish I had this when I was in my PhD program.
As a glimpse into the very tip of the iceberg, this diagram is https://tex.stackexchange.com/a/158740/ generated with 100% Latex code.
Emergent Mind works by checking social media for arXiv paper mentions (HackerNews, Reddit, X, YouTube, and GitHub), then ranks the papers based on how much social media activity there has been and how long since the paper was published (similar to how HN and Reddit work, except using social media activity, not upvotes, for the ranking). Then, for each paper, it summarizes it using GPT-4, links to the social media discussions, paper references, and related papers.
It's a fairly new site and I haven't shared it much yet. Would love any feedback or requests you all have for improving it.
I just extract the titles and look for their respective ids.
The real challenge was how to do that at scale. Only in CS there are well over half a million papers
Also, soon-ish I'm going to add the ability for users to follow specific authors, so you can get notified when they publish new papers.
If you could do it, this would be a dream. My original intent was to be able to look through only papers citing a popular one and filtering the results for ones having at least one author with a set minimum h-index. Using Google Scholar data required using SerpAPI, which has some annoying limitations.
The core goal is obviously just not to miss out on a paper that will very likely be influential while not having to comb through the mountain of irrelevant papers.
What's funny is that Microsoft Academic was the best suited, but was retired in 2021.
Would be nice if I could change timeframe. Top this week, month, year, all time.
Love the concept though. Added it to my Home Screen on iOS
I might add comments down the road if there's enough interest and if there's enough traffic to warrant it. Don't want to add them just yet and have zero comments on everything and it look like a ghost town.
Keep the suggestions coming though as you use it more: matt@emergentmind.com.
is there a site that lists and rates the various LLM models of hugginface.co alongside their various applications?
https://opensource.adobe.com/dc-acrobat-sdk-docs/pdfstandard...
if you look at sections 14.6 through 14.10 you will find quite baroque facilities for representing the structure of documents in great detail, making documents with accessibility data, making documents that can reflow with HTML, etc. Note to mention the 14.11 stuff which addresses problems with high end printing (say you want to make litho plates for a book.)
For that matter sections 14.4 and 14.5 describe facilities that can be used to add additional private data to PDF files for particular applications. For instance Adobe Illustrator's files are PDF files with some extra private data, and https://en.wikipedia.org/wiki/GeoPDF
I like to complain that PDF has no facility to draw a circle but instead makes you approximate a circle with (accursed) Bézier curves but other than that the main complaint people make about PDF is that it is too complicated not that it is lacking this feature or that feature.
Contrast that to a highly opinionated document format like DjVu
https://en.wikipedia.org/wiki/DjVu
which came out around the same time as PDF and is specialized for the problem of scanned documents and works by decomposing the document into three layers, one of which is a bilevel layer intended to represent text. All three layers have specialized coding schemes, the text layer in particular tries to identify that every copy of (say) the letter "e" or the character "漢" is the same and reuse s the same bitmap for them.
There is a reason that most people still use docx for forms even pdf technically support forms.
PS: pdf reader of firefox and chrome don't really supports forms until very late versions.