Show HN: Convert screenshots of equations to LaTeX
mathpix.com
mathpix.com
https://tex.stackexchange.com/questions/1443/what-is-the-sta...
Under API... you're already doing handwriting? This is uh, nontrivial work to say the least. Really impressive.
The endorsements are a nice touch. :)
Made me really curious how far the system goes, what cases break it.
Oh... nevermind. You have a PDF of examples here: https://docs.mathpix.com
It's honing in on equations without getting distracted by nearby Hanzi or Cyrillic, or even pictures of dogs. Wow.
I keep going back to dig through your resources and getting more impressed.
EDIT: I guess my only constructive criticism is that you should brag more. I like a simple landing page, but I think you've earned a short list of examples of corner cases you tackle well, if the whole API is packed into that free app, because they're really impressive.
If I wanted to use this in an application, I'd definitely want to see some accuracy figures on validation data as well as a few failure cases to see whether the output remains reasonable even when it is wrong.
Digital pen input contains more info than the resulting bitmap; strokes are lost while rasterizing.
That info was the reason how old devices were able to reliably recognize characters written by a stylus. It worked well even on prehistoric hardware, such as 16MHz CPU + 128 kB RAM in the first Palm PDA.
Offline is way harder than online.
One note I should make: it was not entirely clear (to me) upon a cursory view of the website, that the purpose of mathpix was to convert handwritten text into LaTeX. For some reason (maybe my coffee hasn't kicked in yet) I thought this was strictly intended to take screenshots of equations on an existing pdf document or a website etc and that will be converted to LaTeX.
My thought at that point was "I wonder if they could do this for handwritten text" and then I looked at the docs and facepalmed..
For reference:
screenshot
(verb) an image of the data displayed on the screen of a computer or mobile device.Also, your definition describes a noun but claims it's a verb.
I also meant noun, not verb. Thanks.
This would be great for blind people, as pdfed latex is extremely non-accessable, and I have to email authors of papers to get the original latex from them, which is often lost.
Best case I could copy and paste paragraphs at a time from a PDF of the textbook (with copy protection removed). Worst case I was retyping or fixing every few words in a sentence.
I was working on this from about 2011-2013. Advances with image processing and machine learning have been significant since then, so there may be much better software available now.
If anyone has ideas or packages they'd recommend, I'd be interested.
[1] http://www.inftyproject.org/en/software.html#InftyReader
Many years ago I "translated" course materials into a form which was accessible to a blind grad student. It was a really interesting job and taught me a lot about accessibility.
I was effectively doing latex, but without all the leading \ characters. It made learning latex comparatively easy.
What interface do you use to read equations? Screen reader speaking the straight latex, or do you have some Middleware to make it more digestible when listened to?
Mathpix only does the equation OCR part.
I've worked on this (for a PDF to HTML application), mail is in profile if you're interested.
https://www.springfieldspringfield.co.uk/view_episode_script...
What kind of sorcery is this!?
Is this using deep learning or "regular" OpenCV or similar?
I would assume it's a highly tuned deep learning algo, but I'm not knowledgeable enough to distinguish a deep learning algo from a pile of rocks...
Edit: Aha, someone already asked this and got an answer.
Great software otherwise
Bug report: it appears that multiline summation subscripts are not recognized correctly. For example, Eq. 8 of [1]. These are often created using \substack as part of amsmath.
Awesome tool!
I strongly suggest you talk to the publishers about integrating your tech into their TeX.
http://lstm.seas.harvard.edu/latex/
Here's how to do it with OpenNMT/PyTorch:
Also it would be nice of some info on the process. Does work entirely locally, or is images uploaded to the cloud?
- \mathcal letters (recognized as non-mathcal)
- long equations (not recognized at all)
- multi-line equations (not recognized at all)
The screenshot2latex tool: https://github.com/rmst/screenshot2latex/blob/master/scripts...
Coming from a grad student who hates writing equations in latex. I will probably try this out.
\left\{ \begin{array} ... \end{array} \right
instead of \begin{cases} ... \end{cases}
?There are actually a number of ongoing research projects to establish standards of semantical mathematical representations. Probably one of the best funded running projects (budget ~10MEUR) which has a work package on this topic is http://opendreamkit.org/ . Work is going on at https://mathhub.info/ from my knowledge. I would like to provide a deep link but the site seems to be in a broken state. Apparently people are working on it right in the moment.
Are these projects aiming at something like what I’m describing? Or are they more about something else like verifying proofs?
On the other hand, there are these research projects which however seem to concentrate on standards rather then actually accumulating semantic knowledge.
I would love to see an adoption of hypertext and semantic mathematical notation in scientific papers. Instead of writing $E=m c^2$ in a (LaTeX) paper, we would instead define the symbols machine readably with a code like
set E = physics/Energy
set m = physics/Mass
set c = physics/constants/speed-of-light
I have never seen actual scientific papers which do this kind of stuff, i.e. which are machine readable.The verbalization if most LaTeX commands can help learn to read the equations. Sometimes.
Thank you so much!
> Mathpix's AI definitely passes THIS Turing test!
I can type an equation into LaTeX more quickly that I can photograph it and then go in and manually correct all the spacing issues. And there are spacing issues. The examples PDF has things that just look horrible. No small spacing or negative spacing to space out things like matrices and integrals etc. If I'm going to manually tweak it anyway I might as well do it manually from the start. Typing it up is something that gets quicker with practice like everything else.
But what I can actually see happening is people not tweaking the output manually. Either train yourself to use TeX properly, or let someone do it for you.
Typesetting does not always need to be absolutely perfect, sometimes it just needs to work.
For someone who isn't 100% fluent in TeX, making small edits is much easier than having to look up all the syntax and symbol names required.
Trouble is I'm not sure if these are absolute rules. I'm sure you can find pathological cases where the spacing will be wrong. Maybe the best thing would be to allow the user to modify the spacing easily after the automatic version is made.
I trust you guys have read the TeXBook? Knuth dedicates two chapters to subtleties of mathematical typesetting.
I will let you do it for me. I usually work late at night just before the deadline. What’s your phone number?
As.usual, it depends on the use case. If you are producing a textbook or something that requires high quality typesetting, you will probably pay for it.
If not, you can use a tool like this one but you will have to accept certain margin of error.
The same complaint could be raised about using a biro (vs calligraphy), using a printing press (vs handwriting) or using a wysiwyg editor (vs the old word processing paradigm).
In my view, this sort of accessibility is a great advance.
Calligraphy is not handwriting. It's an art form and very slow to execute.
The printing press isn't just an "easier handwriting". It's actually harder to use. But it's better. It's closer to calligraphy than handwriting, in fact.
I'm not sure what your argument about WYSIWYG text editors is. In my experience they produce absolutely horrible results because, ironically, they require the user to know more about typesetting than TeX does if high quality output is desired.