I'm willing to bet that if you just treated each frame as an image, it would result in some weird stuff when you played them as a movie.
> penny per frame
Where did this come from?
> penny per frame
Where did this come from?
If you wanted to do this at high res, you would definitely use a latent diffusion model. The autoencoder is almost free to run, and reduces the dimensionality of high res images significantly, which makes it a lot cheaper to run the autoregressive diffusion model for multiple steps.