The results are better than anything that I've ever seen.
What's the catch? Large processing times? Are the results cherry picked? Or what?
I guess it only works for video to video, but that's still amazing!
The results are better than anything that I've ever seen.
What's the catch? Large processing times? Are the results cherry picked? Or what?
I guess it only works for video to video, but that's still amazing!
To my knowledge, the only open source solution that works well for text to video is Zeroscope v2 XL, and v3 is coming soon. v2 is already on par with RunwayML's Gen-2 while v3 is better.
Runway outputs the best video quality and options for video length whilst pika delivers better fidelity to an input image as inspiration. All of this subject to change without notice
The original frames have different wheel angles so simple text prompted img2img frame by frame approach would preserve the motion, but at the cost of interframe consistency.
Here you get consistent look of the scene and no rapid transitions, but the wheel motion is gone.