This doesn't follow for me. There are what, Dozens or Erdos tier problems that got solved with no progress for decades? How does that factor in to your view?
IIUC the argument is that while unsolved there was much work done on them that shows up in the training data. The idea being that the LLM is limited to a small amount of inference over externally supplied data.
And yet the best humans could not use that same available data to solve the problems.