There is no information in this article on how modern LLMs like GPT-6 Astra or Claude Opus 5.5 perform at this game. I have a hard time believing they are bad at it.
No, it doesn't make sense. GPT-6 Astra solves ARC-AGI-3, which consists of a large number of small games which the LLM doesn't know and which it has to solve on the fly with a limited number of turns.