perhaps I'm missing something. Why not start the learning at a later state?
If you sat down to solve a problem you’ve never seen before you wouldn’t even know what a valid “later state” looking like.
This is precisely how RL worked for learning Atari games: you don't start with the game halfway solved and then claim the AI solved the end-to-end problem on its own.
The goal in these scenarios is for the machine to solve the problem with no prior information.
Indeed, this is a key to teaching people to know how to advance. Do not focus on a side, but learn to advance a layer.