That means the questions went over the fence to OpenAI.
I'm quite certain they are aware of that, and it would be pretty foolish not to take advantage of at least knowing what the questions are.
Needless to say, it doesn't bring us any closer to AGI.
The only solution I see here is people crafting their own, private benchmarks that the big players don't care about enough to train on. That, at least, gives you a clearer view of the field.
I'm being completely serious. You are correct, despite the downvotes, that this could not be pushing us towards AGI because if the dataset is leaked you can't claim the G-- generalizability.
The point of the benchmark is to lead is to believe that this is a substantial breakthrough. But a reasonable person would be forced to conclude that the results are misleading to due to optimizing around the training data.