Playing hard exploration games by watching YouTube
arxiv.org
arxiv.org
The AI is rewarded if at each checkpoint the state vector its produced is sufficiently aligned with the videos.
I guess that's the initial training to deal with sparse rewards.
Also interesting assumption to say "harder = fewer rewards". Probably doesn't always apply but is a good generalization.