You can imagine this as a spectrum. On the one end you have models that, at each output pixel, try to predict pixels that are locally similar to ones in previous frame; on the other end, you could imagine models that "parse" the initial input image to understand the scene - objects (buildings, doors, people, etc.) and their relationships, and separately, the style with which they're painted, and use that to extrapolate further frames[0]. The latter would obviously fare better, remaining stylistically consistent for longer.
(This model claims to be of the second kind.)
The way I see it: a human could do it[1], so there's no reason an ML model wouldn't be able to.
--
[0] - Brute-force approach: 1) "style-untransfer" the input, i.e. style-transfer to some common style, e.g. photorealistic or sketch, 2) extrapolate the style-untransfered image, and 3) style-transfer result back using original input as style reference. Feels like it should work somewhat okay-ish; wonder if anyone tried that.
[1] - And the hard part wouldn't be extrapolating the scene, but rather keeping the style.