I think the bigger question is would it be stable enough. Many SD like models struggle with consistency across multiple images (i.e. frames) even when content doesn't change much. Would he a cool problem to see tackled.
Tbh most cost effective would be a conditional GAN though
Then train the model on movies that are color and then turn them black and white.
That way you can train temporal coherence.