This has been reinforcement learning for decades. People say "OMG! The agent is acting and learning and getting better all on its own!" And they're right, in the beginning. Eventually the agent plateaus or collapses.
Supervised learning is much more stable. GTP is supervised learning. Once you start letting the agent choose or modify its own training data, then you're moving towards the much less stable realms of reinforcement learning.