> As we have seen, leading edge LLMs, such as the GPT-4o, can solve very complex math problems.
No… they can’t. That’s like saying a search engine can solve math problems — which it can, in a sense.
I suspect that the people repeatedly saying this simply lack the knowledge to know what really constitutes a ‘complex math problem’.
And of course any half-decent new model can answer this particular question correctly; the designers aren’t stupid or unaware of what the expectations and common traps are. The model itself probably will be able to talk about why testing on such comparisons would be interesting (because it ‘knows’ about how this being a recent meme).