But it takes -- what, two or three sentences? -- to explain how the ghosts should move. The interesting (and important) thing is not that the model gets it wrong at first, it's how easy it is to correct it.
One-shotting something like Pac-Man doesn't prove much. At the end of the day, one-shot fidelity is going to scale more or less linearly with model size/world knowledge. Why wouldn't it?