This OP claims the publicly available models all failed to get Bronze.
OpenAI tweet claims there is an unreleased model that can get Gold.
>we note that the vast majority of its answers simply stated the final answer without additional justification
While the reasoning steps are obviously important for judging human participant answers, none of the current big-game providers disclose their actual reasoning tokens. So unless they got direct internal access to these models from the big companies (which seems highly unlikely), this might be yet another failed study designed to (of which we have seen several in recent months, even by serious parties).
OpenAI likely had unlimited tokens, and evaluated "best of N attempts."
We'll never know how many GPUs and other assistance (like custom code paths) this model got.
Meanwhile high schoolers get a piece of paper and 4.5 hours.
[1] https://chess.stackexchange.com/questions/9959/did-deep-blue...
[2] https://nautil.us/why-the-chess-computer-deep-blue-played-li...