Tbh most cost effective would be a conditional GAN though
Then train the model on movies that are color and then turn them black and white.
That way you can train temporal coherence.
24 frames per second * 60 seconds per minute * 90 minute movie length = 129600 frames
If you could get cost to a penny per frame, about $13k? But I'd bet you could easily get it an order of magnitude less in terms of cost. So $1500 or so?
And that's assuming you do 100% of frames and don't have any clever tricks there.
> penny per frame
Where did this come from?
If you wanted to do this at high res, you would definitely use a latent diffusion model. The autoencoder is almost free to run, and reduces the dimensionality of high res images significantly, which makes it a lot cheaper to run the autoregressive diffusion model for multiple steps.