> It makes more sense to assume that ML models not fitted to a particular game would be bad at it than otherwise.
No, it doesn't make sense. GPT-6 Astra solves ARC-AGI-3, which consists of a large number of small games which the LLM doesn't know and which it has to solve on the fly with a limited number of turns.