A Web UI for Stable Diffusion
github.com
github.com
I will say that this one has had REALLY active development as new features have been coming out, and is pretty polished at this point (albeit I'm using it more as a toy than anything, but it's awesome to have a quick way to use the new features that have been shipping out).
ERROR: for stable-diffusion Cannot start service stable-diffusion: failed to create shim: OCI runtime create failed: container_linux.go:380: starting container process caused: process_linux.go:545: container init caused: Running hook #0:: error running hook: exit status 1, stdout: , stderr: nvidia-container-cli: initialization error: load library failed: libnvidia-ml.so.1: cannot open shared object file: no such file or directory: unknownI tried using a docker container and it took 3 min to generate a prompt. However it seems that 2:45 min is somehow spent on tje GPU and finally the remaining 15 seconds the GPU gets utilized.
I haven't had the time to look into this yet, but it does seem to work.
To run it elsewhere in the cloud, grab a GPU (spot) instance and SSH in.
And what do you do after you SSHed in? The installation instructions seem to be for windows users (click here, then click there ...) is there a linux script that does the installation automatically?
Parent comment alludes to docker-compose up.
https://docs.aws.amazon.com/AWSEC2/latest/WindowsGuide/accel...
With these models it works exactly the same way. Someone dropped millions of rocks and created a formula of unbelievable complexity and what they now did is they released that formula with all their calculated parameters into the world. What you do when you ultimately use Stable Diffusion is you just calculate the result of this formula and that is your image. You never have to process those images.
When programming, it will often take a long time and a lot of code to get to a few final lines that do what you want. You cannot say the final result is a "thumbnail" of all previous efforts. Rather, it is the apotheosis of it.
Some artists spend decades developing a style that looks like a kid could do it as well. Still, there is something unique in there, that a trained eye will recognize. Converting that particular style to a formula and making that freely available is at least somewhat morally ambiguous.
Sure, it's not copyright infringement, but you could argue that this takes away from the hardship the original artist had to go through to perfect their style.
Which was precisely what the discussion was about.
> But you could argue that this takes away from the hardship the original artist had to go through to perfect their style
You could argue the same things about photoshop, a lot of other digital tools, drum machines, the photograph and the phonograph.
Fair use might work but maybe not? If I were to argue against it, I'd probably compare something like a recording of music vs. a MIDI file. Same raw data scaling.
So in a way, hundreds of billions of possible images are all stored in the model (each a vector in multidimensional latent space) and turned into pixels on demand (drived by the language model that knows how to turn words into a vector in this space)
As it’s deterministic (given the exact same request parameters, random seed included, you get the exact same image) it’s a form of compression (or at least encoding decoding) too: I could send you the parameters for 1 million images that you would be able to recreate on your side, just as a relatively small text file.
For any input image? Or do you mean an image generated by the model?
No idea how true that is, but on my windows machine, same params/seed is definitely deterministic.
[0] a help string in the SD source code recommends the ddim_eta parameter (which isn't exposed in most web UI or GUI's, including the OP github) stay at the default 0.8 for deterministic sampling. I have no idea if this means changing the value from 0.8 produces non-deterministic results with the same hardware/os/params/seed. Or if they just mean changing this from 0.8 will make your SD not match the online model but still be deterministic itself. But in my testing, changing this value gives no useful changes to the image generation, so I keep it at 0.8
After running for a while, the adversarial network outputs a seed, and you now have a few characters representing a reasonable approximation of your image.
A kazillion images are used in training, but training consists of using those images to tune on the order of ~5 GB of weights and that is the entire size of the final model. Those images are never stored anywhere else and are discarded immediately after being used to tune the model. Those 5 GB generate all the images we see.
For StableDiffusion, the current model is ~4GB, which is downloaded the first time you run the model. These 4GB encode all the information that the model requires to derive your images.
It's not a search engine, it's self-contained and the closest analogy is that it's a very very knowledgable and skilled artist.
The model (presumably some kind of convolutional neural network) has many layers, every layer has some set of nodes, and every node has a weight, which is just some coefficient. The weights are 'learned' during the model training where the model takes in the data you mention and evaluates the output. This typically happens on a super beefy computer and can take a long time for a model like this. As images are evaluated the output gets better the weights get adjusted accordingly.
Now we as the user just need the model and the weights!
The goal is usually to have the more fundamental parts of your model already working and you thus need way less domain specific data.
Here, you're not training anything, you're running the models (both the CLIP language model and the unet) in feedforward. That's just deploying your model, not transfer learning.
There's a (very ironically named) "Manual installation" section which might seem to be the answer for Linux, but then it's not immediately obvious which preceding sections are Linux or Windows without doing critical thinking.
Stable Diffusion is built on PyTorch. PyTorch mainly has been designed to work with Nvidia cards. However PyTorch added support for something called RocM like a year ago that adds compatibility with newer AMD cards.
Unfortunately RocM doesn't support slightly older AMD cards in conjunction with intel processors.
So my 32gb pretty powerful 2020 16in MacBook Pro isn't capable of running Stable Diffusion.
Any native app will likely have to rely on a remote cloud gpu. And boy, those are fucking expensive. Been researching what I need to stand up a service the last few days and it isn't cost friendly.
And not just any old M1 Mac. Last week I got it running on my 2021 8GB M1 MacBook Air and it's slow. Images at 512x512 with 10 steps take between 7 and 10 minutes to generate.
It's the only thing I do that hits performance limitations on the 8GB machine so there's no regrets on that score, but with the way this stuff is progressing 16GB+ is a realistic minimum for comfortable use.
I've just been following the steps here with default settings: https://replicate.com/blog/run-stable-diffusion-on-m1-mac, but maybe there's a better way to run it at this point?
It's an electron app and you can either download it with the weights, or without and add them separately.
Using that it just took slightly over 5 minutes on my 8GB so it's a little bit quicker for me. Maybe the code has improved since I cloned stuff over a week ago, or maybe it's just different system resources when it is run. Either way, it looks like the easy way I've been waiting for.
Edit: It is however missing the CFG setting.
Unless you want to train the model, Lambda Labs is somewhat cheap:
I plan to wrap things up and put out the source this weekend.
There are a few bugs to iron out before it's ready for prime time. For now, create the folder `~/Desktop/charl-e/samples/` manually before you run it.
A commenter mentioned today it might be possible to pre-download the model and load it into the browser from the local filesystem rather than include such a gigantic blob as an accompanying dependency, fighting different caching RFC's, security/usage restrictions, and anything else that might inadvertently trigger a re-download.
- activate advanced: create prompt matrix and use
@a painting of a (forest|desert|swamp|island|plains) painted by (claude monet|greg rutkowski|thomas kinkade)
- add different relative weights for words in a prompt:
watercolor :0.5 painting :0.2 by picasso :0.3
- Generate much larger images with your limited vram by using optimized versions of attention.py and model.py
https://github.com/sd-webui/stable-diffusion-webui/discussio...
- Generate "Loab the AI haunting woman" if you can (Try using textual inversion with negatively weighted prompts)
https://www.cnet.com/science/what-is-loab-the-haunting-ai-ar...
https://github.com/sd-webui/stable-diffusion-webui/wiki/Inst...
- add RealESRGAN for better upscaling
https://github.com/sd-webui/stable-diffusion-webui/wiki/Inst...
- add LDSR for crazy good upscaling (for 10x the processing time)
https://github.com/sd-webui/stable-diffusion-webui/wiki/Inst...
What's with the "hyper resolution", "4K, detailed" adjectives which are thrown left and right, while we are at it?
One thing that is very powerful with Stable Diffusion is using text inversion ( https://textual-inversion.github.io/ ) - you can add additional input samples to further extend the possibilities beyond what is included in the original model.
You can always run it on a CPU and utilize your RAM instead if needed, though the training might extend to 24+ hours that way.
Edit: Here's an example of someone successfully using textual inversion - https://www.reddit.com/r/StableDiffusion/comments/wz88lg/i_g...
For some genuinely incredible results try this pattern for instruction:
Portrait of {Name of some type of identity such as "Faerie Princess" or "Dragon Queen"} {Name of a celebrity such as "Scarlett Johansson"}, beautiful face, symmetrical face, tone mapped, intricate, elegant, highly detailed, digital painting, artstation, concept art, smooth, sharp focus, illustration, art by artgerm and Greg Rutkowski and Alphonse Mucha and Boris Vallejo and Johannes Voss and Aleksi Briclot and Michael Komarck.
Run several iterations of the same query as some results will have anomalies.
Edit: Got home and was able to double check. It's actually a solid 10 seconds per image with the following settings: seed:466520488 width:512 height:512 steps:50 cfg_scale:7.5 sampler:k_lms. Still quick enough for some fun, but could be annoying if you're need to do multiple iterations a minute.
But yeah the next generation of models would probably capitalize on more memory somehow.
The original txt2img and img2img scripts are a bit wonky and not all of the samplers work, but as long as you stick to dream.py and use a working sampler, I have had good luck with k_lms, then it works great and runs way faster than the cpu version.
Works great on 32gb ram but I'm honestly tempted to sell this one and get a 64gb model once the m2 pros come around. This is capable of eating up all the ram you can throw at it to do multiple pictures simultaneously.
It helps if you consider it all as effectively advanced compression. Everything the model can do is limited by its architecture, the number of parameters in the model, and the accuracy and size of the training data.
The underlying architecture is a transformer (e.g. GPT3) wired to a (denoising) diffusion model.
Current flaws with this approach:
- Transformers seem to approach a "bag-of-words" model, often ignoring the ordering of the words. Among other things, this means that text-to-image models are very bad at "binding attributes" [0]. This is why "a boy wearing a red shirt and a girl wearing a black jacket" may fail (putting the colors on the wrong items, for instance).
- Autoregressive transformers have no means to correct early mistakes.
- Training data is captioned images and the captions are likely noisy and under-specified. Every time it sees a face labeled as "face" - it tries to generate a face from the distribution of _all_ faces in the data. The same goes for the dice. If the dice are just labeled "dice", but don't have a description of how they landed - the model has to guess which angle you're referring to. As a sibling comment points out, this is exacerbated by the relative frequencies of examples of the data in the dataset.
Well, I kid a bit. I’ve seen it produce some amazing results, but, generally, it has a hard time with that. Often faces end up looking blurry or having these creepy, dead white eyes. Hands likewise often end up malformed (seven fingers anyone?) and twisty. But, it seems to have a much easier time generating passable faces in close ups with the right key words. Especially if you give it an input image that already has a clear one. It also seems to have an easier time doing faces it already knows like a celebrity, presumably because it’s using a strong existing influence instead of inventing/hallucinating it.
Supposedly this is improved in their new 1.5 version which is in beta. The software is so compelling that I suspect this will be improved quite quickly. Also, I think either way workarounds will emerge, either by composing with other networks/software (some UIs have GANs for face correction) or the old fashioned way by photoshopping over the blemishes.
> GFPGAN Face Correction: Automatically correct distorted faces with a built-in GFPGAN option, fixes them in less than half a second
So apparently there is still an issue with faces.
Example image directly from Stable Diffusion:
https://i.imgur.com/XSk8fIv.png
And here is that image run through GFPGAN:
https://i.imgur.com/I53AGmh.png
Interesting to note how specialised GFPGAN is, as some of the other details (flowers, hair) seem to be worse in the processed image. I plan to finish this image by manually blending the best of both pictures.
In the US, AI generated art cannot be copyrighted.
Edit:
Some additional details.
The US also denied a copyright for one where the creator listed themselves, with the AI just being a co-creator.
https://www.reddit.com/r/COPYRIGHT/comments/vshypc/the_us_co... (Original article is paywalled, reddit post contains the relevant bits)
In particular: >“Even though you argue that there is some human creative input present in the work that is distinct from RAGHAV’s contribution, this human authorship cannot be distinguished or separated from the final work produced by the computer program,” the office stated.
The US does seem to be a bit of an outlier here. The above work was granted copyright in Canada and India.
In the EU, AI generated artwork is likely copyrightable: https://link.springer.com/article/10.1007/s40319-021-01115-0
The same for the UK: https://www.kilburnstrode.com/knowledge/ai/ai-musings/respon...
Edit2: I'm not a lawyer, this isn't legal advice, go contact one if you actually need legal advice here.
"Because copyright law as codified in the 1976 Act requires human authorship, the Work cannot be registered."
The actual ruling (and a similar USPTO discussions) are about AI generated art and talk extensively about it in the broad case. The stance of these organizations is that AI generated art is not copyrightable. I don't disagree that the line is blurred when you discuss content aware fill, where the AI is working on a portion of it, but the current use of SD, even img2img and multiple prompts, etc., quite clearly falls outside of human authorship as recognized by the US Copyright and Patent offices.
https://www.copyright.gov/rulings-filings/review-board/docs/... https://www.uspto.gov/sites/default/files/documents/USPTO_AI...
Might this change in the future? Possibly. But as it stands today, I would not make any plans that assume you can secure the copyright (in the US) to anything made with SD.
Edit: Going through and noting that I'm not a lawyer and this isn't legal advice, don't listen to some random on the internet for legal advice, get a lawyer if you need it.
I think these are answering a slightly different question, as they are asking if the AI itself can hold the copyright on the output. A bit like if someone tried to copyright an image and assign “Photoshop” as the author.
The question above is maybe closer to asking if the person using an ML model can get copyright on the output, in that case there is a person trying to own the copyright, so I suspect it would not be rejected so easily.
Who is?
The original question I replied to: >If someone runs the model on their own hardware, do they "own" the images generated?
This seems to be straightforward - Thaler tried to receive the copyright for the artwork generated by his Creativity Machine. He was denied, because the copyright office does not believe that a neural network generated image has human authorship.
From the Copyright office paper: "he [Thaler] was “seeking to register this computer-generated work as a work-for-hire to the owner of the Creativity Machine.”"
>A bit like if someone tried to copyright an image and assign “Photoshop” as the author.
This is also clearly outside of the scope of copyrightable work per the reasoning given by the copyright office.
Both questions are thoroughly answered at this moment unless Thaler wins his appeal.
Edit: Going through and noting that I'm not a lawyer and this isn't legal advice, don't listen to some random on the internet for legal advice, get a lawyer if you need it.
> Who is?
Thaler is. I’ve only read the intro sections of the documents you linked to, so I may have missed something more fundamental later, but the key points seem to be:
> The author of the Work was identified as the “Creativity Machine,” ... the Work “was autonomously created by a computer algorithm running on a machine”
and:
> Thaler must either provide evidence that the Work is the product of human authorship or convince the Office to depart from a century of copyright jurisprudence. He has done neither.
So in this case they are asking if the AI can be the author.
Whereas the question in this thread was:
> If someone runs the model on their own hardware, do they "own" the images generated?
In that case, a human is providing a prompt to the model (providing creative input), and asking if they themselves count as the author (a human rather than a neural net), so it seems like a significantly different case.
>In that case, a human is providing a prompt to the model (providing creative input), and asking if they themselves count as the author (a human rather than a neural net), so it seems like a significantly different case.
I don't know that I specifically agree with this, but this is probably due to me having read additional articles on similar filings, including one where someone took a photograph, applied a style transfer AI to it, and then tried to copyright the resulting image, and was denied, because the copyright office found that there was not evidence that the work was a product of human authorship.
Andres Guadamuz (a lawyer specializing in IP law, senior lecturer at Sussex university, and a proponent of AI generated work being copyrightable) discusses a lot of this in https://www.technollama.co.uk/dall%c2%b7e-goes-commercial-bu... - but the most relevant part to this discussion is "For the most part, the legal consensus appears to be that the images do not have any copyright whatsoever, and that they’re all in the public domain."
The user experience for DALL-E, StableDiffusion, Midjourney, etc. are all essentially the same - craft a prompt, fine-tune it, get artwork out, so his discussion should be broadly applicable to all of these similar tools.
I happen to be in the UK, and this happens to match my expectations, but it does strongly imply more regional variation than I’d have guessed:
> The situation may be different in the UK, where copyright law allows copyright on a computer-generated work, the author of which is the person who made the arrangements necessary for the work to be created. This, in my opinion, is the user, as we come up with the prompt and initiate the creation of the specific work. I think that there may be a good case to be made that I own the images I create in the UK.
This was a Style Transfer AI - it takes a source image and recreates it in the style of a painter.
In this case, the person both took the photo that the style was transferred to, and selected the style and a variety of variables. The US Copyright office still felt that his contribution was not distinguishable from the work that the AI did.
I'll note that this is very US specific - there are a lot of counter-examples of other countries allowing for the copyright of work like this, including the EU, UK, Canada, India, etc.
Nobody would ever think about Adobe or Nikon having copyright claims over your pictures. For me it's just a tool, the artistic part is providing a good description/base image, refining and choosing the best output.
Anyway, I'm not a lawyer and we probably live in different countries, so it'll be interesting to wait for the first lawsuit.
But it is illegal to share pictures of Eiffel Tower, for example.
People do it, but they shouldn't.
If I put a picture of Eiffel Tower at night in a book or any other kind of commercial product, I have to pay to use it. Doesn't matter that it's there for my eyes to see it.
The question is: are the images generated by an hyper accelerated learning machine using copyrighted material without the author's consent legal?
I think they shouldn't be and the data included in the training should be free or licensed.
Grab that, install Docker, install Nvidia's Docker integration, copy the example Docker-env file, and docker-compose up is all you need.
Edit: here's a gist with exact steps I used: https://gist.github.com/geerlingguy/384ed4aba35e3118f2a0f358...
First they came for illustrators, then they came for UI designers.