You could compose multiple network, one with the goal of learning how to interact with the world through the screenshots. This would probably require a lot of frames of the game from many diverse locations and scenarios--as the network would need to learn how battles works, items (though technically not strictly necessary), and general dialogue and menu interactions.
Then other networks could then try to learn to encode different intermediary goals trained on a bunch of noisy runs that complete that goal. Collecting this data seems tedious and difficult though and is a whole project in itself.
So my intuition is yes, but I don't think simply making my current model bigger would get us there. I see you have experience with world models--what would you try? ;)