The redditor was using img2img to do style transfer frame by frame, which is why we have the jumping faces all the time (“instability”).
This isn’t a limitation of neural nets, as early as 2018 we’ve had stable style transfer for videos, see https://medium.com/element-ai-research-lab/stabilizing-neura...