See also this recent blog post + paper by OpenAI, they designed curiosity in an agent by having it test how well it can predict the output of a randomly initialized neural network from the game's current frame + input. The better the prediction, the more its mental model of the world is correct, and the less it gets rewarded (encouraging it to go and find unexplored situations). https://blog.openai.com/reinforcement-learning-with-predicti...