1) Train a model to predict what happens next given an input 2) Each frame, predict the next frames given all possible inputs 3) Choose the input that maximizes uncertainty
I would expect this to learn to avoid deaths relatively quickly. It doesn’t need to be good at knowing what will happen next, just better at recognizing specific dead ends (e.g. spikes or holes).
"AI researchers have typically tried to get around the issues posed by by Montezuma’s Revenge and Pitfall! by instructing reinforcement-learning algorithms to explore randomly at times, while adding rewards for exploration—what’s known as “intrinsic motivation.”
But the Uber researchers believe this fails to capture an important aspect of human curiosity. “We hypothesize that a major weakness of current intrinsic motivation algorithms is detachment,” they write. “Wherein the algorithms forget about promising areas they have visited, meaning they do not return to them to see if they lead to new states."