Goodness me this is the stuff of nightmares!
I asked it for "A dog catching a treat in slow motion, it's chops flapping around comically"
https://imgur.com/a/dog-catching-treat-slow-motion-its-chops...
I asked it for "A dog catching a treat in slow motion, it's chops flapping around comically"
https://imgur.com/a/dog-catching-treat-slow-motion-its-chops...
https://imgur.com/a/warriors-who-are-cows-fighting-with-edo-...
Somewhere between "base compute" and "4x compute".
So maybe you "just" need to know how to create a certain type of diffusion transformer model and then train on a ton of videos, but with an adequate amount of compute. Which is probably a LOT for training and inference to get more realistic results.