Stable Diffusion: Is Video Coming Soon?
metaphysic.ai
metaphysic.ai
This will be a huge game changer when it occurs. Whether it be for deep fake videos, creating custom content, or making a new season of your favorite tv show that was cancelled too early. The possibilities are endless.
This is probably not in the near future (i.e. this year), but I doubt it is very far off.
You also could put a language model on top of your prompting system. So "gandolff kicking ass" gets translated into " Page XXX, Paragraph XX from LOTR "
It's only a few frames, but they are entirely generated from text - no seed image or interpolation required.
I'm sure these types of capabilities will come at some point, but no current model can do it. It requires more than just projecting motion into a scene.
It will take a lot of compute to compile the script and render the video.
What it would need to be able to do to get from here to there is understand some concepts. The first being "characters". On reddit there was beautiful image that recently won first place in an art contest and its quite frustrated some of the art community. When I was looking at it I thought it was awesome, but wondered at the ability to create another hundred or so images in that same 'world' that the created image was showing. I would want to do something like give it the prompt "tired old medieval knight with a mace and shield" and have it create the character then be able to name it "Tom" or something and feed it more prompts for that characters like "Tom is sitting in a forest brooding" and have it create the same exact character but in a different context.
That would be pretty game changing for opening up amature web comics to a large body of people who have ideas and tell stories but have no art skills to speak of - my stick characters are crooked :(
Another trick to approach this problems is specifying the random seed, this will cause the same image being generated by the same prompt without any randomness. When you now change the prompt you get an image that is very similar to the first one, but with the variation included. Somebody used that to age a woman across 100 years[2] with quite stunning results. Even works with gender or style changes.
[1] https://textual-inversion.github.io/
[2] https://www.reddit.com/r/StableDiffusion/comments/wq6t5z/por...
[3] https://www.reddit.com/r/StableDiffusion/comments/wq6t5z/por...
Last week PhilFTW explained "How To Create a Complete Graphic Novel in ONE Day" with Midjourney in a YouTube video [1]. He uses five tools:
- Midjourney (to generate images)
- InferKit (to generate the story text)
- Word (to rearrange the story text to fit into some narrative)
- Comic Life 3 for iPad (to place the images and text in comic book panels)
- Affinity Designer (to design the cover and export everything to print, Kindle, and Blurb)
Err, no? As you can tell by the name itself, the 5B dataset is larger.
> "Unlike autoencoder-based deepfake content, or the human recreations that can be achieved by Neural Radiance Fields (NeRF) and Generative Adversarial Networks (GANs), diffusion-based systems learn to generate images by adding noise"
This is confused. Diffusion is orthogonal to NeRF. For example, here is a paper that uses both: https://arxiv.org/abs/2112.12390
> "Within days of release, the open sourced Stable Diffusion code and weights were packaged into a free Windows executable"
That's not how it became popular.
> "Additionally, at the time of writing, Google Research has just released a similar system called DreamBooth, which likewise ‘tokenizes’ a desired element into a distinct ‘object’."
The approaches are actually very different. DreamBooth uses fine-tuning, unlike textual inversion.
Also, diffusion models learn to remove noise (guided by a description of the undiffused image), not add it.
Though, it seems you don't have to use noise as the information destruction mechanism; blurring works, and I wonder if there isn't something better for animation.
It seems closer to what Mac OS X Leopard's Photo Booth was able to do 15 years ago[1], than to a "Stable Diffusion for Video".
[1] - https://web.archive.org/web/20071018033504/https://www.apple...
https://www.tiktok.com/@karenxcheng/video/713806710521107588...
Even a small model in early training learns to do proper shadowing and lighting and mostly makes mistakes in the 3D high-level object space, like wrong number of wheels or legs or fingers and stuff like that.
I don’t think they do, but I can’t prove it, either way.
> (just like humans by the way).
Oh, I don’t think humans have it either (mostly), since humans don’t have 3D vision; humans have the equivalent of wiggle stereoscopy¹, which gives some hints of depth, and humans have the additional advantage of the time dimension. People do have some intellectual capacity to reason about 3-dimensional shapes, and some people can even rotate things in their heads with ease. Blind people might have it too, since their concept of the world was not created by this pseudo-3D visual input. But mostly, people don’t think in 3D.
I think we can see this by looking at drawings by children. Child drawings are dominated by the concepts important to children: Faces, hands, etc. Concepts, not actual images. And it’s all in 2D, as was the majority of art for much of history.
The sensory input modalities and specifics do not control or limit the internal representations; a sufficiently capable neural network will extract the most efficient rep to predict the input data, which is moving 3D objects.
That it's difficult to "render" this to 2D by painting is not very surprising.
Well, some people reportedly can, but it’s rare. You say you can imagine walking around your small city, and infer that you have the city conceptualized as a 3D object, and can rotate it at will, in any direction. But how do we know that you simply have not memorized your admittedly small city? Can you imagine what going around the city upside down (i.e. like walking on your hands) would look like, with ease? It should be just a simple rotation, right?
https://openai.com/blog/dall-e-introducing-outpainting/
The image might look consistent at first sight, but if you look closer, the dimensions are all over the place. They reflect the method by which they were created: reproducing observed patterns without deeper understanding.
At some low level, pixel patterns are rendered and I guess you could say that this is "reproducing observed patterns". Would you say that about a 3D game engine as well that does the same when it textures local regions of pixels? A network is layered for this particular reason. The lowest layer will have less "understanding" than the higher layers.
we are going all the way to the holodeck (my guess is latent diffusion on weights, aka hypernetwork, of a 4d NERF)
Still the Transframer results are interesting and has encouraged me to look into it
Today still the AI misses local context as the images are more of trained from Open images annotated in English by experts. But imagine if you have single image annotated by multiple people in different languages then what happens to AI capabilities
Though I'd venture that the first "novel to feature film" or "novel to TV series" algorithm won't just be an upscaling of this tech....
So cool for abstract art but not for storytelling or following a script. Unless you are OK with the content being visually inconsistent like an acid trip.
Smaller image models had the same problems with logical inconsistency just because they didn't have sufficient general understanding of how visual concepts.
The same is almost certainly true of video - early smaller models will likely create janky movements/motion, however once they've seen enough video to understand how a person walks, how a scene is framed etc.. there's no reason we couldn't get to the same level of maturity as today's image models.
I think the real issue will come from labelling - most video is only going to be labelled simply with basic info/captions without detailed descriptions of the camera pan, movement of subjects. The amount of text required to accurately describe a scene is much larger than a still image and I'm not sure how once would go about collecting this.
Anyone aware of any other open source projects that have a known list of requested features and way for community to express support for them, for example bounties?
Something else which would be possible is the use of a model like SD in combination with a frame interpolation model like [2] as a video generator. Use SD to generate key frames, feed these to FILM and let it generate the intermediate frames and you should get video.
Below is a mini tut where the content from Dall-e it piped to Ebsynth and then DAIN. https://www.instagram.com/reel/Ch7aV2mjWOD/
In summary: Dall-e generated the outfits, Ebsynth mapped those outfits to a range of frames (instead of having a new random artwork on each frame) & Dain smoothed the transitions between each outfit change.
Another example of this concept is here: https://www.instagram.com/reel/ChmyFNoDHZY/
https://replicate.com/andreasjansson/stable-diffusion-animat...
The idea is that it takes two txt2img prompts and animates between them. The result isn't 100% there yet, but I think this idea has some legs on the 2D side of things because of the way it's able to keep a consistent composition. (One of the core issues with txt2img is that while it can produce nice images, it can product just one - animation, or even a series of story boards requires a lot of manual intervention from the creative.
The fact that no one has been able to demonstrate it convincingly makes me think that it might actually be quite hard.
Presumably you'd want to be director and be able to control camera angles. You might want to have a cartoon or 3D render visual style. Unreal engine prompt tag even more literal. Most important, you will want actors. Anything alive will have to move convincingly and intentionally. Some might even bring up the ethics of allowing the movie to end.
Before proper video, I think we'll first have to see tools for music, 3D rendering and animation that bring down difficulty by orders of magnitude.
There have been some actual video explorations of SD's latent space, this example amazes me: https://twitter.com/karpathy/status/1559343616270557184
Like, feed in some text and pose questions, like people have done with GPT-3.
You can try TextSynth and see the result is nowhere near as good.
Hope you don't actually think that's what these models do.