What have you to say about methods such as these:
"LeanDojo: Theorem Proving with Retrieval-Augmented Language Models" (2023) https://arxiv.org/abs/2306.15626 https://leandojo.org/
https://news.ycombinator.com/from?site=leandojo.org
( https://westurner.github.io/hnlog/#story-38435908 Ctrl-F "TheoremQA" (to find the citation and its references in my local personal knowledgebase HTML document with my comments archived in it), manually )
"TheoremQA: A Theorem-driven [STEM] Question Answering dataset" (2023) https://github.com/wenhuchen/TheoremQA#leaderboard (they check LLM accuracy with Wolfram Mathematica)
"Large language models as simulated economic agents (2022) [pdf]" https://news.ycombinator.com/item?id=34385880 :
> Can any LLM do n-body gravity? What does it say when it doesn't know; doesn't have confidence in estimates?
From https://news.ycombinator.com/item?id=38354679 :
> "LLMs cannot find reasoning errors, but can correct them" (2023) https://news.ycombinator.com/item?id=38353285
> "Misalignment and Deception by an autonomous stock trading LLM agent" https://news.ycombinator.com/item?id=38353880#38354486
That being said, guess and check and then develop a fitting casual explanation is or is not the standard practice of science, so
https://news.ycombinator.com/item?id=38124505 https://westurner.github.io/hnlog/#comment-38124505 :
> What does Mathlib have for SetTheory, ZFC, NFU, and HoTT?
> Do any existing CAS systems have configurable axioms?
Does LeanDojo have configurable axioms?
From https://news.ycombinator.com/item?id=38527844 :
> TIL i^4x == e^2iπx ... But SymPy says it isn't so (as the equality relation automated test assertion fails); and GeoGebra plots it as a unit circle and a line, but SymPy doesn't have a plot_complex() function to compare just one tool's output with another.