Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
arxiv.org
arxiv.org
This is exactly the problem that needs to be solved. The yes-man nature of LLMs is the biggest inhibitor to progress, as a model that cannot self evaluate well cannot learn.
If we solve this though, combined with reasoning, I feel somewhat confident we will be able to achieve “AGI,” at least over text-accessible domains.