It is only a limitation of the interface that we're interacting with. There is no reason it couldn't backpropagate towards a better solution when told that it's wrong. OpenAI probably aren't letting it train online lest some jokers try to teach it racism and other bullshit etc.