Animating Prompts with Stable Diffusion
replicate.com
replicate.com
The possibilities for non-skeuomorphic media format is pretty insane, especially once we get into animation territory.
I've been using Midjourney and SD to create a sci-fi filmverse called SALT, here's some details about how I put together, including all available "episodes": https://twitter.com/fabianstelzer/status/1565085199322456069
The generated art concept is pretty cool but why couldn't we leave crypto out of it for once?
That said, SALT will be 100% open to anyone, regardless of whether they’re using crypto or not
You say "a fun idea in crypto gaming is composability", but I don't see why that is specific to cryptocurrency.
(FWIW, I think the idea of composable game worlds created by ML prompts is absolutely fascinating, and I also think a future in which all money is outside the control of the state and corporations is interesting. I just don't see the relationship between the two).
I'm yet to see anything from this world made by the latest AI generated images boom.
For example I really really like Midjourney, it creates images that feel artistic and all but I start to think that d I'm misjudging it because it appears to be a that tool makes great combinations that look fascinating because they are so novel and out of this world.
Crystals growing over the electronics, porcelain bubbles, viola that turns into plasma etc... all amazing but all these are a genre in art. Combining things, making things transition into other things, making them look like something else - all that are procedures that humans can master(it's just that the computer can do it much more quickly).
Considering that all this is simply teaching a computer to predict stuff by degrading images and trying to re-create back again, I think this is going to make a revolution in tooling when we can actually guide the output precisely. Right now it's just fascinating toy for "out of this world" image creation and anything made by AI looks like the imaginary bridges and buildings that you can find on Euro banknotes(made like that in order not to favour particular country over others). That said, I think the toy in its current state has high explorational value.
Have you not just described creativity (and genetic recombination) in the abstract? I understand this ability to combine and synthesize and cross-breed to be one of the more important components of human intelligence, from which almost everything else wonderful about us is derived. And it shouldn't escape not to replication and recombination is also the centre of biological information processes of life.
Something that recombines things better and more "creatively" than us in it's very nascent days, that feels quite important to me. Respectfully, it doesn't feel like simply a "genre of art" that it's doing better than us
Humans are not an island. I recommend watching the “everything is a remix” series it does a good job of expounding on this
> I'm describing the nature of creativity and no, the creative process is not simply fitting things together somehow.
Many non-naive folks in the sciences would agree that creativity (as far as the universe is concerned about human activity as a physical phenomenon) is basically just aimless recombination of semiotic structures, seeking stability that helps it (and associate structures) to persist. It's not unlike the function that genetic recombination performs in the biological strata of information. Genetic recombination are creativity in the biological substrate. Our version is just much more highly dimensional, but it's the same shit (the same forces in the universe) underlying it.
Your saying human creativity is "something more" feels to me like an example of elevating the subjective conscious experience of it
In these animations it's easy to see it, as the shape recognition is not very stable. For example, in the last video when the prompt changes from "people" to "bears" you can see a backpack turning into a bear head, which then turns into an open-mouth bear head. And then, the arm of the person carring the backpack is also turned into more bear heads.
The next steps in the evolution of this technique should be in exploring the relation between noise and subjects, so that you can create variations of the same image maintaining the recognizable parts stable.
This is some amazing work that takes advantage of it: https://twitter.com/xsteenbrugge/status/1558508866463219712?...
It cuts off before the full montage but you should get the idea.
The zoom effect is scaling the original frame larger while keeping the frame, then performing image to image generation on it, or using the newly defined image in the frame and sending it back through diffusion (at least that’s my guess).
It’s a different process than what DALLE has since inpainting does not overwrite already generated pieces. Stable diffusion can also do inpainting.
You can make sliding images this way. Slowly translating the image out of the frame and then filling in the blank space with inpainting.
The linked technique zooms and doesn't use masks, hence the instability.
Could anyone explain why this phrases is repeated everywhere?
So it might be that all images have a distinctive look, or influences of it, and this phrase is becoming kind of an inside joke/meme.
It's a stupid temporary problem
It's by deforum. Here are links to their Discord, GitHub, and Colab: https://deforum.github.io/
I'm guessing it's running out of memory perhaps.
What I would really love is to have multiple prompts. So, for example, prompt 1 is a fast moving description of the action. Prompt 2 is a slow moving fade between artistic style.
> Provide 'frame number : prompt at this frame', separate different prompts with '|'. Make sure the frame number does not exceed the max_frames.
Checking out the code, in wouldn't be too difficult to add a system for movement keyframes either.
Some good examples of this being done with Dalle2 last month -- https://youtu.be/TW2w-z0UtQU?t=244
rot_mat = cv2.getRotationMatrix2D(center, angle, scale) # the zoom variable is passed as scale
https://github.com/deforum/stable-diffusion/blob/5241ce95058...Edit: having looked carefully at the video, these are pretty different. Inpaining keeps the original crop the same, whereas this version allows the model to reinterpret the original.
(I know… but let me think it, guys)
I don't have a fast machine myself, but I would not mind renting a VM somewhere to play with it.
Any tips?
There are quite a few tutorials for getting started with SD.[2] They tend to explain how to install it on your CPU, but some also explain how to build your own instance on Google Colab . I know this one in Spanish[2], and you can look for more. [3]
[1] https://colab.research.google.com/github/altryne/sd-webui-co...
[2] https://youtu.be/5z223SxlAcA?t=1910
[3] https://www.youtube.com/results?search_query=stable+diffusio...
And be able to chain together my own workflow, combining different tools.
So I would prefer to rent a VM and not use Google Colab.
It looks like they've got a nice python notebook: https://github.com/deforum/stable-diffusion/blob/main/Deforu...
For other cases I would reccomend this repo which has a user script feature: https://github.com/AUTOMATIC1111/stable-diffusion-webui#user...
https://replicate.com/andreasjansson/stable-diffusion-animat...