I think a reasonable assumption is that there is an interaction between an LLM, a https://en.wikipedia.org/wiki/Computer_algebra_system tool, a human prompting with deep math expertise, and lots of compute that explains hitting upon the remarkable cancellation.
The original tweet implied that the whole thing was done while the author was watching the World Cup final.
I know it’s tempting to hope that a human did the “real” work here, but if some special insight was put into prompting, the author kept it to himself, and there is no reason why they would hide this since it would elevate their own status.
Have encountered a similar flavor in programming, wrote it off until I saw someone point out how garbage in garbage out they tend to be. If you hand any frontier model dogshit and ask it to do something simply, the result is often not great.
But! If you spend 20 minutes having it comb through and clean up with something like jscpd, then tell it to step through with a debugger, gather profiling traces, etc... very likely it will yield meaningful improvements or catch some corner cases. If it doesn't, anyone with experience is going to tell it to try something else, or that it isn't good enough, as opposed to accepting the first result.
You can recreate this by disabling web search and asking a model about the conjecture and then giving it his post. I've tried a few and their initial responses range from "this is a meme I'm not even going to verify it" to vaguely insulting chains of thought, concerns about the need to be careful because you're clearly nuts or stupid, then falling back on remedial explanations. After a few nudges they all eventually work through it, accept it, and apologize.
IMO its reasonable to imagine a situation where someone is having a beer or two watching The Big Game, asking an LLM to do something stupid for fun and landing somewhere like this on the magic jump to conclusions mat.
It is premature to assume the author is not going to share more information in the future about the mathematical insights to narrow down the search space for this counterexample.
I think it would also explain their opacity towards the process. Being able to solve such well known problems in a nice replicable 1-2-3 way would be far more effective marketing than their complete opacity outside of the result, which suggests that they feel transparency is not in their best interest for some reason.
And what do you get from working at the company? Likely a rather massive token/processing budget. The companies opacity towards the path to these discoveries also makes this further probable as 'spend millions of dollars in tokens' is a somewhat less attractive narrative than the implied narrative of 'just use Fable.'
Well, mathematicians not working for Anthropic/OpenAI are heavily disincentivised from reporting that their discoveries were made using AI. If e.g. the idea that resolved the Mahler conjecture came from AI, it's not like we'd ever know.
Was Anthropic doing this earlier this year? with some of Smale's open problems? Did those improvement go into training the LLM that was used a week ago? Seems like a plausible factor.
We're definitely still in the computer chess phase.