Solving some formal math olympiad problems
openai.com
openai.com
To be fair, however, these proofs are using several of the automatic proving methods built into Lean, like linarith, nlinarith, field_simp, and so on. Each of these are hundreds of lines of manually tweaked heuristics.
It's still quite interesting. I think if we had a good integration of AI methods into the Lean runtime itself the theorem prover could be excellent. You could have a cooperative model where the human mathematician writes down statements, and the AI fleshes out a proof or catches typos that you make. Right now Lean has similar functionality but the automatic proving isn't as good as it could be.
I just wish the Lean community wasn't stuck on the Lean 3 -> Lean 4 migration for the past year. I'm getting flashbacks of Python 2 -> Python 3 ....
For 5/6 of the problems, Z3 solves them automatically and instantly, with the same problem encoding as the problems in the post. (Problem 2 involves factorial, so that one can't be straightforwardly translated to Z3.)
Code: https://gist.github.com/anishathalye/0d5cd359adcde85fac6bf76...
"We achieved a new state-of-the-art (41.2% vs 29.3%) on the miniF2F benchmark, a challenging collection of high-school olympiad problems." [2]
This system's performance is already better than the average scores of students who take these tests, who are usually very good at math.
The test/validation sets are drawn from AMC, AIME, and IMO which are competitive math competitions at the high school level (IMO being the hardest). Even though the topics are generally limited to high school curriculum, solving the problems often requires a degree of creative thinking and knowledge synthesis such that many/most college graduates cannot solve them.
1. artists (deep dream, etc.)
2. cab/truck drivers (self driving cars)
3. call center operators (robocalls)
4. programmers (copilot)
5. mathematicians
who is next?
We’re just waiting for hardware engineers to produce faster hardware/software engineers to make the implementation more efficient. ;-)
Having said that, the problems solved here are fairly simple. I can do half of them in my head (how old are the students doing these quizzes?). Getting from the written form to a description that Lean can work with probably is harder than getting Lean to solve them. Skimming the paper, it doesn’t look like they included that first step.
For example, determining the 3d description of a photographed scene (a difficult problem) you can just enumerate all the possible scenes, raytrace them, and pick the most probably one (e.g. the one with the simplest description).
It would be interesting to eval copilot on the same benchmark as I'm pretty sure it can close some of the proofs still.