Pix2tex: Using a ViT to convert images of equations into LaTeX code
github.com
github.com
I think this paper was the first one to do OCR on LaTeX: http://cs231n.stanford.edu/reports/2017/pdfs/815.pdf The paper describes an Encoder-Decoder architecture with CNN encoder and LSTM based decoder.
Some recent cool work he's been doing: https://www.youtube.com/watch?v=lx1XcTdhalU.
https://rohankulkarni.me/files/notes/heidelberg_qft/12_2.pdf ("12.2 Diagrammatic expansion of partition function for Yukawa theory")
Another example is standard model of particle physics. There’s a way to write down the Lagrangian of the standard model very compactly: https://visit.cern/content/standard_model_formula_t_shirt
But if you expanded all the implied sums and terms, it probably would be monstrous. After all it is supposed to contain all the terms that this figure contains for example: https://commons.wikimedia.org/wiki/File:Standard_Model_of_El...
Also, ask string theorists to show you calculations that they need to find the largest paper to perform. I’ve heard people doing calculations where a single line is the width of an A1 paper.
The language is important. The more you “understand” something, the simpler it is (over simplification here.) eg I would not consider the Einstein equation expanded out to be natural after understanding it. (Well Maxwell’s equation can be summarized in 1 single elegant equation as well.)
But the reason I’d consider that optics formula to be monstrous is that no matter how you group it into smaller pieces (ie refactoring it), there’s no way to hide the fact that it is not elegant at all. There’s no “understanding” there, it is just so happen the exact solution looks like that.
To put it that way then, often fundamental “master equation” are simple in some ways, but exact solutions to some particular manifestation of that master equation is often quite ugly and monstrous. In that sense then rather than quoting Einstein equation I’d quote its solution eg one with mass and spin and electric charge.
P.S. the solution of quartic equation is also a good example of this category
I've gotten a lot of requests to do whole equations, but I feel that would massively increase the complication of the app for not that much benefit? How often do people want to convert a whole bunch of equations into LaTeX? My use case is usually writing my own equations and forgetting the command for a specific symbol, or looking for a symbol that looks something like X.
The longer I live, the more I'm interested in saving all of my data into text files that I can parse later without vendor lock-in concern. Maybe other open formats as well, best tool for the job, ya know.
It's just a zip file containing a bunch of XML. And the slides XML isn't beautiful/super nice but not super ugly either. Naively processing it is lossy, but not as lossy as converting it to text.
And most images end up as png in them. The most annoying thing is images with data (like equations).
> latex2sympy parses LaTeX and generates SymPy symbolic CAS Python code (w/ ANTLR) and is now merged in SymPy core but you must install ANTLR before because it's an optional dependency. Then, sympy.lambdify will compile a symbolic expression for use with TODO JAX, TensorFlow, PyTorch,.
mamba install -c conda-forge sympy antlr # pytorch tensorflow jax # jupyterlab jupyter_console
https://news.ycombinator.com/item?id=36159017 : sympy.utilities.lambdify.lambdify() , sympytorch, sympy2jaxThere are a number of ways to generate tests for functions and methods with and without parameter and return types.
Property-based testing is one way to auto-generate test cases.
Property testing: https://en.wikipedia.org/wiki/Property_testing
awesome-python-testing#property-based-testing: https://github.com/cleder/awesome-python-testing#property-ba...
https://github.com/HypothesisWorks/hypothesis :
> Hypothesis is a family of testing libraries which let you write tests parametrized by a source of examples. A Hypothesis implementation then generates simple and comprehensible examples that make your tests fail. This simplifies writing your tests and makes them more powerful at the same time, by letting software automate the boring bits and do them to a higher standard than a human would, freeing you to focus on the higher level test logic.
> This sort of testing is often called "property-based testing", and the most widely known implementation of the concept is the Haskell library QuickCheck, but Hypothesis differs significantly from QuickCheck and is designed to fit idiomatically and easily into existing styles of testing that you are used to, with absolutely no familiarity with Haskell or functional programming needed.
Fuzzing is another way to auto-generate tests and test cases; by testing combinations of function parameters as a traversal through a combinatorial graph.
Fuzzing: https://en.wikipedia.org/wiki/Fuzzing
Google/atheris is based on libFuzzer: https://github.com/google/atheris
Clusterfuzz supports libFuzzer and APFL: https://google.github.io/clusterfuzz/setting-up-fuzzing/libf...
What would be really nice is, if I could feed slides.pdf to something like this, and it did OCR on every handwritten text (english or equation), and put the output as an invisible layer under the text. Will make the slides searchable.
I understand though, OCR on handwritten equations, is a very difficult problem.
This is very much idle daydreaming. When you write out a big equation on the wall, what happens then? Does it need to be validated? Or does it go directly into a paper and no computation is performed upon/with it?
Also, for a lot of people with addiction to think on blackboard where you don't have to worry about anything else. It is easy. Take a photo of what you did, erase the board, write new things and take another photo. And when you are done and want to preserve this in some notes or copy to paper, just use an equation recognition tool and your life is much easier.
It is a productivity tool that saves and efforts, it will not be going to make you a super researcher/scientist.
Usually, whoever does the derivation, or someone who wants to understand things properly, will do computations on multiple steps of the derivation from the start to the finish. A lot of these computations can be done by hand - you don't need a computer. A lot of computations should be done by hand - even if they could be done by a computer - because you only get a feel for the equations if you play with them with your hands. To quote Dirac, 'I consider that I understand an equation when I can predict the properties of its solutions, without actually solving it.' That comes from solving a lot of them by hand.
Yes, oftentimes, doing numerical or symbolic computation with a computer helps. But is the pain point of that having to type the equation into the computer. Hardly. It would be nice, but nothing ground breaking.
I have heard about "mixture of experts" as being a potentially important advance, and also of course about multimodality. So I found this: https://github.com/YeonwooSung/LIMoE-pytorch
Now show us a version that takes into account actual representations and errors, produces an optimal implementation of the calculation for accuracy, and explains why it is. :)
I fed the equation image (screenshot at the right frame from their gif then cropped) into ChatGPT (GPT4-V) and it correctly deciphered the equation and gave the correct LaText code.
Why was the repo removed?
It might be more error prone though, didn’t test it extensively