I'm pretty sure such upscaling systems train against ground truth higher res video and purposefully downres'd versions of that video. So, at the very least, we'd need to first get higher res video of the same phenomena once. Though maybe you could style transfer from frames that have been improved by hand.