Neat, I don't understand what they mean by having embedded a reward video into the set. Is that a video where copying the behaviour will deliver victory?
The AI is rewarded if at each checkpoint the state vector its produced is sufficiently aligned with the videos.
I guess that's the initial training to deal with sparse rewards.