With these models it works exactly the same way. Someone dropped millions of rocks and created a formula of unbelievable complexity and what they now did is they released that formula with all their calculated parameters into the world. What you do when you ultimately use Stable Diffusion is you just calculate the result of this formula and that is your image. You never have to process those images.
When programming, it will often take a long time and a lot of code to get to a few final lines that do what you want. You cannot say the final result is a "thumbnail" of all previous efforts. Rather, it is the apotheosis of it.
Some artists spend decades developing a style that looks like a kid could do it as well. Still, there is something unique in there, that a trained eye will recognize. Converting that particular style to a formula and making that freely available is at least somewhat morally ambiguous.
Sure, it's not copyright infringement, but you could argue that this takes away from the hardship the original artist had to go through to perfect their style.
Which was precisely what the discussion was about.
> But you could argue that this takes away from the hardship the original artist had to go through to perfect their style
You could argue the same things about photoshop, a lot of other digital tools, drum machines, the photograph and the phonograph.
Fair use might work but maybe not? If I were to argue against it, I'd probably compare something like a recording of music vs. a MIDI file. Same raw data scaling.
So in a way, hundreds of billions of possible images are all stored in the model (each a vector in multidimensional latent space) and turned into pixels on demand (drived by the language model that knows how to turn words into a vector in this space)
As it’s deterministic (given the exact same request parameters, random seed included, you get the exact same image) it’s a form of compression (or at least encoding decoding) too: I could send you the parameters for 1 million images that you would be able to recreate on your side, just as a relatively small text file.
For any input image? Or do you mean an image generated by the model?
No idea how true that is, but on my windows machine, same params/seed is definitely deterministic.
[0] a help string in the SD source code recommends the ddim_eta parameter (which isn't exposed in most web UI or GUI's, including the OP github) stay at the default 0.8 for deterministic sampling. I have no idea if this means changing the value from 0.8 produces non-deterministic results with the same hardware/os/params/seed. Or if they just mean changing this from 0.8 will make your SD not match the online model but still be deterministic itself. But in my testing, changing this value gives no useful changes to the image generation, so I keep it at 0.8
After running for a while, the adversarial network outputs a seed, and you now have a few characters representing a reasonable approximation of your image.
A kazillion images are used in training, but training consists of using those images to tune on the order of ~5 GB of weights and that is the entire size of the final model. Those images are never stored anywhere else and are discarded immediately after being used to tune the model. Those 5 GB generate all the images we see.
For StableDiffusion, the current model is ~4GB, which is downloaded the first time you run the model. These 4GB encode all the information that the model requires to derive your images.
It's not a search engine, it's self-contained and the closest analogy is that it's a very very knowledgable and skilled artist.
The model (presumably some kind of convolutional neural network) has many layers, every layer has some set of nodes, and every node has a weight, which is just some coefficient. The weights are 'learned' during the model training where the model takes in the data you mention and evaluates the output. This typically happens on a super beefy computer and can take a long time for a model like this. As images are evaluated the output gets better the weights get adjusted accordingly.
Now we as the user just need the model and the weights!
The goal is usually to have the more fundamental parts of your model already working and you thus need way less domain specific data.
Here, you're not training anything, you're running the models (both the CLIP language model and the unet) in feedforward. That's just deploying your model, not transfer learning.