Most of this doesn't discredit your overall point, though.
Community maintained spreadsheet of the runs: https://docs.google.com/spreadsheets/d/e/2PACX-1vQDvsy5Dt_-P...
What NVIDIA has here is a generic "evolution" harness, which can be used for any problem.
I think it would be fair game to allow OpenClaw, Hermes, Codex, Grok Bot, this NVIDIA thing, to compete, as long as they don't have ARC-AGI specific skills, toolset.
True as that may be, it may be better to optimize models for some amount of memory versus forcing some token count based on a reasoning level, right?
GPT-4 was decidedly not capable of beating Pokemon 18 months ago. I doubt it would be able to complete a single level. I don't think people realize how large the advances in model capabilities have been. GPT-4 in a modern harness is absolutely horrendous compared to modern models.
Have you ever played pokemon?