Stablevideo: Text-driven consistency-aware diffusion video editing
rese1f.github.io
rese1f.github.io
"We present the content deformation field (CoDeF) as a new type of video representation, which consists of a canonical content field aggregating the static contents in the entire video and a temporal deformation field recording the transformations from the canonical image (i.e., rendered from the canonical content field) to each individual frame along the time axis. Given a target video, these two fields are jointly optimized to reconstruct it through a carefully tailored rendering pipeline. We advisedly introduce some regularizations into the optimization process, urging the canonical content field to inherit semantics (e.g., the object shape) from the video."
The results also look very stable/impressive.
1. The boat video transforms the coast into rolling waves and looks super weird
2. The swan/duck video looks better than the others, but the lighting is obviously wrong when looking closer. Looks like a cardboard cutout bird.
3. The car video looks like a video game from the 2000s, with low quality textures and wheels not turning.
Again, super interesting to see how it can come up with these "by itself", but utterly useless at the moment.
I wouldn't call it entirely useless though - it makes for an interesting surreal effect. I could see something really cool coming out of this in the A Scanner Dark movie vein.
I particularly liked Rusty Car in the Desert.
(See stable diffusion/llama/ chatgpt)
There will be businesses that actually make money on these technologies, and they will be research heavy (even a 5% improvement is a big deal) as things are still getting figured out.
I could see speed dropping back towards 2017 like rates, but I kind of doubt we will ever see an true ai winter like the 90's early 00's
The field is just too young with too many things as of yet untried, along with the fact that I doubt funding will dry up any time soon. (There are too many interests, from Nvidia wanting to sell more chips, to Microsoft wanting to sell more productivity, to defence, and political concerns between the US and China.)
Yes, it won't go on forever, but also this time seems qualitatively different from the past AI cycles. (Granted I was not alive then)
The field has been around since the 50s with various summers and winters, with each summer having people saying it's now too big to fail, with ever increasing resources and time being spent on it, only for it to eventually stagnate again for some time. If there is one field in computer science I wouldn't call "too young", it would be AI. The first "true" AI winters happened in the 1970s/1980s, and second one in the late 1980s. You seem to have missed some of them by a large margin.
It's the natural movement of ecosystems that are hyped a lot. They get hyped until there is no more air, and it goes back into "building foundations" mode until someone hits gold and the cycle repeats all over again.
I'm sure things will develop, but develop into flawless midjourney-but-for-video? literally only time will tell, its a fools errand to extrapolate
Since human brains during dreams (lucid or otherwise) can generate coherent scenes, and transform individual elements in a scene, diffusion based models running on cpu/gpus should eventually be able to do the same.
That the human brain is exactly equivalent in function to our current model of a neural network is a huge, unproven hypothesis.
Indeed technically that might not be possible due to probabilistic nature of these models and may require a whole different technology. But one thing for sure is that enough labour and capital is going into it so the chances are not little.
The results are better than anything that I've ever seen.
What's the catch? Large processing times? Are the results cherry picked? Or what?
I guess it only works for video to video, but that's still amazing!
To my knowledge, the only open source solution that works well for text to video is Zeroscope v2 XL, and v3 is coming soon. v2 is already on par with RunwayML's Gen-2 while v3 is better.
Runway outputs the best video quality and options for video length whilst pika delivers better fidelity to an input image as inspiration. All of this subject to change without notice
The original frames have different wheel angles so simple text prompted img2img frame by frame approach would preserve the motion, but at the cost of interframe consistency.
Here you get consistent look of the scene and no rapid transitions, but the wheel motion is gone.
I keep seeing these type of webpages to promote papers but haven't found the template yet.
It seems they forked from somebody else and then changed the content to match their paper.
(1) with enough high-quality training data, «AI» models should be able to output H265 / H266 / AV1 directly, can achieve simplicity and reduce artifacts by skipping an inferior compression step and leveraging temporal elements
(2) if AI video compression (as demoed by nvidia) becomes standard, the training data and generated data will become [more] «AI-native», boosting these efforts by miles
2. as long as we can again port the algo's to dedicated hardware, which are on mobiles a must for energy efficiency for both encode and decode