This harness is really moving the goalpost by defeating the entire point of the test. Instead of seeing the strength of a model's world view, its ability to internally derive and intuit rules, and its ability to keep track of game state over time, we're just letting the AI cheat. This is just the LLM equivalent of running a chess engine to the side.
And this harness would not work in a remotely complex game and relies on the fact that Arc-AGI-3 is a focused test that only made the games as complicated as they needed to be for current model performance.