Stable Diffusion XL 1.0
techcrunch.com
techcrunch.com
I clicked through the links in the article, since they sounded technically interesting. They led to AI-generated porn. Those, in turn, led to pages about training SD to generate porn. Now, two disclaimers:
1) I am not interested in AI-generating porn
2) I haven't followed SD in maybe 6-9 months
With those out-of-the-way, the out-of-the-box tools for fine-tuning SD are impressive, well beyond anything I've seen in the non-porn space, and the progress seems to be entirely driven by the anime porn community:
https://aituts.com/stable-diffusion-lora
10 images is enough to fine-tune. 30-150 is preferred. This takes 15-240 minutes, depending on GPU. I do occasionally use SD for work. If this works for images other than naked and cartoon women, and for normal business graphics, this may dramatically increase the utility of SD in my workflows (at least if I get around to setting it up).
I want my images to have a consistent style. If I'm making icons, I'd like to fine-tune on my baseline icon set. If I'm making slides for a deck, I'd like those to have a consistent color scheme and visual language. Now I can.
Thanks creepy porn dudes!
The other piece: Anyone trying to keep the cat in the bag? It's too late.
Its not entirely driven by porn communities, and the porn communities driving it aren’t entirely anime porn communities (and the anime communities driving it aren’t entirely porn communities.)
But, yeah, the anime + porn/fetish art + furry + rpg art + scifi/fantasy art communities, and particularly the niches in the overlap of two or more of those are, pretty significant.
> If this works for images other than naked and cartoon women
It does, and while it may not be large proportionally compared to the anime-porn stuff, there’s a lot of publicly distributed fine tuned checkpoints, LoRas, etc., demonstrating that it does.
I use it for D&D art generation. I can have a piece of art that somewhat matches every location/scene I have planned. If things don't match my plans I can generate 8 images and pick the best in about 2 minutes. I talk to a lot of other DMs who use it in a similar way.
It's not great with specific details, I plan to commission someone to draw the party when the campaign is over. But for things like a fantasy magic shop with potions, or a fantasy dungeon exterior, or a forest of mushroom trees, it's more than good enough for concept art to throw into Roll20. I couldn't afford 5-10 pieces of custom concept art per game, nor could I come up with the ideas for them 2 hours beforehand and have them ready for the session.
I will choose to intentionally misread that :)
It should appear here at some point, currently only the VAE was added:
(I've been having fun with this for a few days. https://huggingface.co/stabilityai/stable-diffusion-xl-base-... Not sure there's much of a difference with the 1.0 version.)
> The refining process has produced a model that generates more vibrant and accurate colors, with better contrast, lighting, and shadows than its predecessor. The imaging process is also streamlined to deliver quicker results, yielding full 1-megapixel (1024x1024) resolution images in seconds in multiple aspect ratios.
Sounds pretty impressive, and the sample results at the bottom of the page are visually excellent.
A 70B parameter model would be very slow and vram hungry, hence very expensive to run.
Also, image generation is more reliant on tooling surrounding the models than pure text prompting. I dont think even a 300B model would get things quite right through text prompting alone.
In fact, most interesting papers since Imagen show that you get more mileage out of scaling the text encoder part, which is, of course, a Transformer. This is what drives accuracy, text rendering, compositionality, parsing edge cases. In SD 1.5 the text encoder part (CLIP ViT-L/14) takes a measly 123M parameters.[1] In Imagen, it was T5-XXL with 4.6B [2]. I am interested in someone trying to use a really strong encoder baseline – maybe from a UL2-20B – to push this tactic further.
Seeing as you can throw out diffusion altogether and synthesize images with transformers [3], there is no reason to prioritize the diffusion part as such.
1. https://forums.fast.ai/t/stable-diffusion-parameter-budget-a...
"A picture is worth a thousand words" - I wonder how (in)accurate this popular saying turned out to be? :D
That’s actually how this whole party got started. DALL-E (the first one) was a transformer model trained on image tokens from an early VAE (and text tokens ofc). Researchers from CompVis developed VQGAN in response. OpenAI showed improved fidelity with guided diffusion over ImageNet (classes) and subsequently DALLE2 using pixel space diffusion and cascading up sampling. CompVis responded with Latent Diffusion which used diffusion in the latent space of some new VQGANs.
The paper you mention is interesting! They go back to the DALL-E 1 method but train two VQGAN’s for upsampling and increase the parameter count. This is faster, but only faster than originally reported benchmarks using inferior sampling methods for their diffusion. I would be curious if they can beat some of the more recent ones which require as few as 10-20 steps.
They also improve on FID/CLIP scores likely by using more parameters. This might be a memory/time trade off though. I would be curious how much more VRAM their model requires compared to SD, MJ, Kandinsky.
The same goes for using T5-XXL. You’ll win FID score contests but no one will be able to run it without an A100 or TPU pod.
Is this still true in 2023? Sure, back in the dark ages it seemed like a 860M model is just about the limit for a regular consumer, but I don't see why we wouldn't be able to use quantized encoders; and even 30B LLMs run okay on Macbooks now.
If we are talking about Stable Diffusion, the reality is that... more parameters mean it will be hard to run locally. And let me tell you something, the community around Stable Diffusion only cares with NSFW... And want local for that...
Stable Diffusion 2 was totally boycotted by the community because they... banned NSFW from there. They had now to allow it again on SDXL.
Also, more parameters mean it will be more expensive to community finetunners to train as well.
Not affiliated in anyway and not very involved in the space. I just wanted to generate some images a few weeks ago and was looking for somewhere I could do that for free. The link above lets you do that but I suggest you look up prompts because its a lot more involved than I expected.
For those not aware, here's an interesting fact about ArtBot (and the AI Horde in general) -- we've been running an A/B test with Stability.ai for the last 3 weeks or so related to SDXL [1].
Any time a user generates an image using SDXL_beta on the AI Horde, they get two images back. They pick which image they think is best for the given prompt. This data is sent back to Stability.ai in order to help improve their image models.
In a similar vein, LAION partnered with the AI Horde earlier this year in order to gather aesthetics ratings for improving various image datasets. [2]
It's a cool little open source community and there's just a ton of stuff going on.
[1] https://dbzer0.com/blog/stable-diffusion-xl-beta-on-the-ai-h...
I just took the ones I liked and then deleted out the words that were specific to that image and left the ones that were providing the style of the image. So for example on the first one I would delete "an cute kitsune in florest" but would keep "colorfully fantast concept art". Then I just added a comma separated list of the of the features I wanted in my picture. It took a lot more trial and error than I thought and adding sentences seemed to be worse than just individual words. I am sure I barely scratched the surface of interfacing with the tool correctly but the space is moving so fast its not the kind of thing I want to spend my time learning right now just to have that knowledge deprecate in 6 months.
But I wouldn't say it's the "best." Just trained on images that weren't taken from unconsenting artists.
Google: The Last Ben Stable Diffusion Colab
for a way to not run it locally, but get all the features.
This may just be due to the iterative denoising approach a lot of these models take but they only seem to work well when creating raster style images.
In my experience when you ask them to create logos, shirt designs, illustrations, they tend to not work as well and introduce a lot of artifacts, distortions, incorrect spellings etc.
As for directly generating vector images, there's nothing yet. Your best bet is generating vector-looking raster and tracing it.
I can run SDXL 1.0 offline from my home. I can’t do this with Midjourney.
A closed source model that doesn’t have the limitation of running on consumer level GPUs will have certain advantages.
So speed can vary wildly depending on how you're choosing to use it. And that's without even getting into the wide variance of hardware.
But generally speaking, it will usually be significantly faster than one image per minute.
StableDiffusion needs you to be way more specific than Midjourney. MJ will fill in the gaps of your prompt to get a better image. SD usually won't.
MJ photos are higher quality with easier prompting IMO, but with a distinctive style. Even if you ask it to mimic some other style, it has that midjourney feel.
I mainly it for generating setting or character images for a D&D game. I use Midjourney more for characters.
[0] This is at ~25 iterations.
A couple of accordion pictures do look passable at a distance.
Another test: how well does it do at drawing a woman waving a flag?
One thing that strikes me is that it generates four images at a time, but there is little variety. It's a similar looking woman wearing a similar color and style of clothing, a similar street, and a large American flag. (In one case drawn wrong.) I guess if you want variety you have to specify it yourself?
AI models seem to be getting ever better in resolution and at portraits.
It's hard to believe we're only 8 months into this industry, so I imagine we'll start seeing smaller footprints soon.
Gpt3 is 36 months old. Dalle-e is 28 months old. Even StableDiffusion is like 11 months old.
TBH devices just need more ram for coherent output though. Llama 13b and 33b are so much "smarter" and more coherent than 7B with 3 bit quant.
TBH the UIs people run for SD 1.5 are pretty unoptimized.
Someday is today: from the official announcement: “SDXL 1.0 should work effectively on consumer GPUs with 8GB VRAM or readily available cloud instances.” https://stability.ai/blog/stable-diffusion-sdxl-1-announceme...
https://github.com/invoke-ai/InvokeAI
Edit: Spec required from the documentation
You will need one of the following:
An NVIDIA-based graphics card with 4 GB or more VRAM memory. 6-8 GB of VRAM is highly recommended for rendering using the Stable Diffusion XL models
An Apple computer with an M1 chip.
An AMD-based graphics card with 4GB or more VRAM memory (Linux only), 6-8 GB for XL rendering.As an aside, does this irritate anyone else?
"You must have Python 3.9 or 3.10 installed on your machine. Earlier or later versions are not supported. Node.js also needs to be installed along with yarn"
I don't like having to install npm when an existing dev stack (python) is already present.
Regular SD can run in less than 2 GB of VRAM with Easy Diffusion.
1. Installation (no dependencies, python etc): https://github.com/easydiffusion/easydiffusion#installation
2. Enable beta to get access to SDXL: https://github.com/easydiffusion/easydiffusion/wiki/The-beta...
3. Use the "Low" VRAM Usage model in the Settings tab.
Would someone be kind to explain what the current state of the art in image generation is (how does this compare to Midjourney and others)?
How do open source models stack up?
Also what are the most common use cases for image generation?
That has been said, based on the configurations of these models, we are far from saturating what the best model can do. The problem is, FID is terrible metrics to evaluating these models so like LLM, we are a bit clueless about how to evaluate them now.
If you want to get arty, then state of the art for out-of-the-box typing in a prompt and clicking "generate" is probably MidJourney.
But if you're willing to spend some more time playing around with the open-source tooling, community finetunes, model augmentations (LyCORIS, etc), SD is probably going to get you the farthest.
> Also what are the most common use cases for image generation?
By sheer number of image generations? Take a guess...
Cat images right?
Catgirl images, to be precise.
For the longest time I thought it was google imaging things and doing some photoshop to make things look like Pixar because it was so bad.
I've found that I rarely get a usable image completely as-is. It might take 5 or 10 generations to find something sort of ok, and even then I end up erasing the bad parts and letting it in-paint (which again takes multiple attempts). The T-rex had like 7 legs and two jaws, but was otherwise close to what I wanted... just keep erasing extra body parts until the in-painter finally takes a hint.
I was also going to do a few book covers for some Babylon 5 books, but it does so bad on celebrity faces. Looked like Koenig's mutant love child with Ernest Borgnine. Dunno what to do about that. I keep wondering if I shouldn't spend the next 10 years putting together my own training set of fantasy and science fiction art.
The "Comfy" tool is node based and you can string both together which is nice. Although if you aren't confident in your images you don't need the refiner for a bit.
Comfy and A1111 are based around the original SD StabilityAI code, but the implementation must be pretty similar if they could add the base model so quickly.
You get deduplication, easy swapping of stuff like VAEs, faster loading, and less ambiguity about what exactly is inside a monolithic .safetensors file. And this all seems more important since SDXL is so big, and split between two models anyway.
Fine tuners will quickly take it further, if that’s what you’re after.
I seem to remember the issue with 2.x is that they removed all the commercial art from top-notch illustrators from the training data due to the backlash, so it was just way worse at generating great-looking things, which is all the user cares about. So the community stayed on their custom-trained models derived from SD 1.5 (which, yes, often included porn).
What you said also happened, but the main thing was the base model didn't have a great concept of human anatomy. Apparently it was really hard to train for anything else as well.
SDXL is a bigger model. There are some subjective comparison posts with SDXL 0.9, but I can't see them since they are on X :/
And yeah, X is weird to type out too.
And/or hand-specific LoRa and/or a workflow using something like ADetailer extension in A1111 that applies a model to recognize hands [0] and then inpaints them.
[0] recognition models are also provided for people, faces, and eyes, too, and it can use additional custom models for other things.
With SD2.1 I’d generate 100 or so images using inpainting and the “good fingers” was about 1% hit rate. If it’s up to 10% that’d be great because generating 10 images takes just a few seconds on an A100
Isn't the best part of a meal eating after you've not had anything to eat for a while? The best part about a kiss that you've quenched the pain of missing your partner?
The best part of art is that you haven't seen anything good in a while?
Scarcity is an underappreciated gift to us, and the relative scarcity per capita is in a sense what drives us to connect with other people, so that we may be priveleged to witness the occasional spark of creativity from a person, which in turn tells us about that person.
Although that sort of viewpoint has been declining for some time due to the intensely capitalistic squeezing of every sort of human endeavor, AI brings this to a whole new level.
I think if those making this software thought a bit about this, they might second-guess whether it is truly right to release it. Just a thought.
Can you please chisel it on stone tablets for me?
That will really help me appreciate it.
Personally I think this guy has a point and he's pointing to something that I don't believe a lot of ai art advocates have considered: the attention economy. There isn't actually any scarcity in the current market for artistic content. There are literally millions of people producing art every day, and the market is very winner take all. There are few artists who are able to support themselves on their work, and their skills are exceptional and specifically in demand. My theory is that the supply of AI generated content will be so vast, and the perception around it will be that it's low effort and low quality, that so called AI artists are going to have trouble distinguishing themselves in a market where they're saturating it and all using the same models. I think your perception of this market is flawed. Art is not a fungible good like food or clothing.
I wouldn't advocate not growing vegetables. But today we grow them in a monoculture for instant availability everywhere, and those monocultures are susceptible to disease and also are not terribly ecologically friendly. AI is like an ultimate monoculture of diseased fruits.
Also, I believe counter-progressive to be a good thing. Human beings should not progress in certain ways, as we don't have the wisdom to use the technology we have developed.
Humans in general cannot appreciate things very well, and computers and AI will only make it worse.
1. Create scarcity, 2. Flood the market to reap short term gains with market-disrupting technology 3. Creat new scarcity by creating new products
But no, I wouldn't want to hand wash my laundry more often. For the same reason probably I still prefer using a lighter when having a BBQ than a flint.
As for the downvotes, I don't mind. I try and present a critical view of technology, but I expect the downvotes because almost everyone here has the perspective that technology is a tool and progress is generally a good thing, which are two statements that I wholeheartedly reject.