I'm expecting next gen video codecs to all be neutral network based.
Ie. every frame is generated from the previous internal state and a few more bytes of input data which defines how the car moves or who the bad guy shoots.
Current state of the art systems require huge GPU's and minutes per frame. But smaller models will do a decentish job faster, and better hardware and algorithms are coming out every few months.