Large Language Models for Mathematicians (2023)
arxiv.org
arxiv.org
https://xenaproject.wordpress.com/2024/12/22/can-ai-do-maths...
https://xenaproject.wordpress.com/2025/01/20/think-of-a-numb...
These from the mathlib founder are much more informative!
There's a blurb about the history here: https://leanprover-community.github.io/papers/mathlib-paper....
https://huggingface.co/spaces/Qwen/Qwen2-VL
You can also run a smaller model locally if you have enough VRAM, for example Qwen2.5-VL-7B-Instruct:
https://github.com/QwenLM/Qwen2.5-VL?tab=readme-ov-file#usin...
Also works reasonably well with hand-written equations.
For searching through similar equations, you can probably embed each as a high-dimensional vector and then search for the closed vector. Here is a ranking of text embedding networks:
https://huggingface.co/spaces/mteb/leaderboard
Or if you want something more deterministic, parse the LaTeX equation to create an abstract syntax tree for which there are plenty of similarity measures.
So when you injest all the latex, you get the semantics, latex conventions, an variable naming of each school math for free
I wonder if from the observation data of particle physics we have, a "physics" model could "infer" a hard mathemical theory using those previous maths models (and probably other tools) which would fit this very observation data.
ScyFy: if those mathemical theories are beyond us, we would need other models in order to try to extract some predictions which could be interesting for us to verify with some "real-life"/reasonable experiments.
Now it's out of date, automated RL on math problems seems to work and scales with compute. As we scale available compute 100x over the next 5 years and reduce cost of compute by around 10x over the same time frame, it will become increasing clear that LLMs running for a long time are capable of replacing most mathematics research.
Are LLMs training on the AST parses of the symbolic expressions, or token coocurrence? What about training on the relations between code and tests?
Benchmarks for math and physics LLMs: FrontierMath, TheoremQA, Multi SWE-bench: https://news.ycombinator.com/item?id=42097683
https://deepmind.google/discover/blog/ai-solves-imo-problems...
Disclosure: I work at Google, but not on AI.
Just call it LLMs for ML Researchers, but then it doesn't sound anywhere as exciting?
The paper notes GPT 4 can solve it (they seemed to have asked ChatGPT 3.5 - this paper is old by AI standards, the first version being from Dec 2023).