Math Olympiads benchmarks at the time were reliant on training data stuffing / Benchmaxxing and this was the absolutely correct statement at that time in that particular context (Dramatic drop of performance on out of training set questions).
We hadn’t broadly moved into lean proof harnesses and brute force compute spend at the time yet and the statement itself continues to be true for LLMs on their own.