Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
It is described in their methodology: https://arcprize.org/policy
It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.
Which LLMs participate on private set? Open weight LLMs only?