You can make decently accessible PDFs but it's lots of work, you need Acrobat on the producer' side and might also need it on the consumer's side. Free tools don't even come close. There's also the fact that the process of making accessible PDFs in Acrobat isn't itself accessible.
With that said, the way screen readers treat HTML math certainly isn't perfect, it's geared more towards school children than anything above calculus. I'm probably going to stay with my LaTeX source files for now. At least ArXiv offers those, not many sites do. To be fair, that approach also has its own set of problems (particularly when people use some extra fancy formatting in their math equations, making the markup hard to read), but I find this to be the best approach for me so far, at least on AI/ML papers.
Say I have Equation \ref{eq}. Why can't I just say "plot \ref{eq} for x from -6 to 11" and get my graph?
And yes, I know about pgfplots, PSTricks, TikZ etc. But in all those cases, I need to define the same equation twice, in different syntax to boot. It's kind of unsatisfying.
Pretty much for the same reason you cannot press a word and get a pop-up dictionary definition in a paper book.
An often cited example: what is f(x+y) ? Is it function f with x+y as its argument, or constant f multiplied by (x+y) ? TeX gives you no clue.
Or what is this i in your equation? Is it an index variable, or a square root from minus one?
You as a human figure this out by looking at the context and using domain knowledge. So does a "TeX to HTML/MathML converter". It is ultimately built on heuristics, and cannot be otherwise.
That's why I said basically "for the same reason a paper page is not interactive". It was designed this way!
The goal of TeX was to generate beautiful printed page. The need for semantic structure was not anticipated. To do semantics you need a "semantic version of MathML", or a language used by Wolfram's product, etc.
A simple example is ‘\sum’ which provides no way to capture the expression being summed over - because that’s not necessary for typesetting. That’s not the case in, say, MathML.
Writing MathML is no fun though because mathematical formulae are visually ambiguous and we rely on the context to know how to read them, e.g. does ‘f(x - 1)’ mean function f called with argument x - 1, or does it mean variable f multiplied by x - 1?
The amount it scrolled probably depended on the aspect ratio of the window, so it might be multiple key presses to scroll an entire column.
I would assume that the majority of persons on HN are not looking at their keyboard as they type.
I'm not deeply familiar with the state of that art, but it seems like recovering the metadata from a PDF generated by LaTeX would be no more impressive than many other things we're currently seeing language models achieve?
It was not designed to provide semantic information, unfortunately. So getting anything other than visual representation out of it is hard.
They argued that PDF was superior because the publisher could control how it looked and it looked the same everywhere but the point is that it should not. Things such as font size and line spacing should be at the control of the consumer, not the publisher. This isn't simply blind people but for instance also persons with dyslexia who use particular fonts to make it easier to read for them. Or in my case, someone who simply gets a headache from fronts and line-spacing that is too big. I've also been using darkmode everywhere for so long now that reading black text on a white surface on a screen gives me a headache.
https://tex.stackexchange.com/questions/485593/how-to-write-...
This problem is harder than you one would think naively.
The problem with this: you need to create a new standard, get everybody to agree to it, and get busy scientists who are concentrating on content and not representation to adapt this new standard in their writing, essentially requiring them to change their habits and spend extra time on writing (which many of them hate), for no obvious gain from their point of view.
I am not saying it's not possible, or not worth it, but it is not easy and simple either.
Besides, in HTML one can directly link to the relevant part.
Being able to link directly to the relevant part is irrelevant (pardon my pun!). Such links are machine-readable, not human-readable. Scientific text need visual citations and being able to name the referred part for reading comprehension.
And Harvard-style citations (AKA name-date) exist for a reason; when your read a paper even in interactive format it helps when you can recognize citations to certain papers and not having to click on them or memorize numbers.
Other styles have their own advantages and disadvantages; that's why they all exist and used by this or that journal, and no consensus on a single "right" style was ever reached.
[1] Great explanation here https://tex.stackexchange.com/questions/57717/relationship-b...
Anyway, if you (or anyone else reading this) has suggestions I'd really appreciate it!
This seems a massive gap in the market - many institutions have funding earmarked for such things.
What kind of turn-around time would be practical? Could you point me to any typeset mathematical braille that would be an example of a solution to your problem? Is Nemeth the only important standard, or are others important for you too?
I'm wondering if it's practical to set this up as back-office work here in Vietnam. There are some outlying provinces here where there are very few job opportunities. Job opportunities for the blind also round down to zero here (e.g. I could hire for proofreading). Maybe there's room to do something cool here.
Keep in mind that most blind people who speak English fluently but don't live in an English-speaking country (myself included) can't read English braille, or at least not well. Because of how voluminous Braille is, it uses contractions, single characters that replace common words and character combinations like "the", "would", "ing" or "ed". Those tend to be language specific, never taught outside their country or countries of use, and hard to get accessible electronic materials for. The math codes are completely different too, we use something derived from Marburg, while English-speaking countries use Nemeth. Even basic characters like + and - differ between those two, not to mention more complicated structures. It's not just the dot patterns that are different but also the design principles, like where you put spaces or when you can omit "begin fraction" / "end fraction" characters.
What would be very useful for me to be able to typeset myself are small things -- homework, quizzes, and (to a lesser extent) exams. Since homework and quizzes often have to adapt to what I actually covered in class, which may or may not match the syllabus, it's hard to rely on sending this out to be typset by others. (Exams are a little easier since they're usually done days ahead of the actual date.)
AFAIK Nemeth is the only standard that matters. If I can typeset a document, send it to the student, and they can get it on a braille display (no need for this to be on paper), it would solve a ton of problems.
Throughout college, my first question to most of my professors of math subjects was "do you do LaTeX, and can you give me your source code." Most said yes, and that's how we worked. LaTeX in, LaTeX or PDF out, depending on what the professor preferred.
The amount of LaTeX you need for calculus 1 isn't that great, you could probably teach it to a relatively bright student if you had an hour or two to spare, and then give them the source files. If you have the time, I'd suggest producing "stripped" versions of your files, with as little markup as possible to get your point across and no fancy formatting unless absolutely necessary. The amount of hoops some books and papers jump through to "look nice" drives me crazy.
You could also consider producing, teaching and consuming ASCII math, which seems like an even simpler and friendlier format. I couldn't really use it much in my school career for boring technical reasons, but it looks like a promising option.
One of my students was taking chemistry at the same time, which is (I think) much tougher for blind students. But they also had more teaching assistants for the course.
https://www.boia.org/blog/why-justified-or-centered-text-is-...
Having a two-column theme, or left-aligned vs justified themes, could be workable in the long run. I hope that we get to see some browser extensions modding the pages before too long.
The reason for the current justified text is that it is the default aesthetic for a LaTeX-based article, and a lot of authors expect it.