Stable Diffusion 3: Research Paper
stability.ai
stability.ai
There was one thing I was curious about... I skimmed through the executive summary of the paper but couldn't find it. Does Stable Diffusion 3 still use CLIP from Open AI for tokenization and text embeddings? I would naively assume that they would try to improve on this part of the model's architecture to improve adherence to text and image prompts.
1. CLIP-G/14 (OpenCLIP)
2. CLIP-L/14 (OpenAI)
3. T5-v1.1-XXL (Google)
They randomly disable encoders during training, so that when generating images SD3 can use any subset of the 3 encoders. They find that using T5 XXL is important only when generating images from prompts with "either highly detailed descriptions of a scene or larger amounts of written text".
Especially the side of the bus.
Still impressive we've come to this but we're now in the uncanny valley which is a badge of honour tbh.
Dream on is a bit better but looks very messy, badly spaced, inconsistent letter sizes.
Of the people I know in the CG industries using SD in any sort of pipeline, they're all using Comfy because a node based workflow is what they're used to from things like Substance Designer, Houdini, Nuke, Blender, etc.
This is also the reason why the generated images have this characteristic high contrast and saturation. Better models usually need to rely less on CFG to generate coherent images because they fit the training distribution better.
Or did we lose Stable Diffusion to SAAS also? Like we did on many of the LLMs which started of so promising as for self hosting goes
> In early, unoptimized inference tests on consumer hardware our largest SD3 model with 8B parameters fits into the 24GB VRAM of a RTX 4090 and takes 34 seconds to generate an image of resolution 1024x1024 when using 50 sampling steps. Additionally, there will be multiple variations of Stable Diffusion 3 during the initial release, ranging from 800m to 8B parameter models to further eliminate hardware barriers.
I'm not even sure what the use case is.
I'd rather wait 30 seconds and get a much higher quality image than some mediocre image in 1 second.
Hell, even if it took 5 minutes per image, and produced even better images, I would prefer that.
Some use cases might be generating profile pictures or banners for users or unique profile pictures for bots in online games. Discord, steam, social media, whatnot, you could just type what you want your profile picture to be and make it on the fly. They're small, aren't expected to be extremely high quality, and cheap enough.
Testing on https://fastsdxl.ai/ - "high quality profile picture of a cartoon cat holding a Bouquet of flowers"
To be clear it's not perfect, but this is a fairly complex prompt and I find the majority of seeds would be "good enough" for thumbnail profile pictures. I think we're almost there for "cheap good enough" usecases.
Places like steam, discord, etc you very rarely see profile pictures above that size.
It is just art. I think AI art shows what obsessed gadget makers for profit we have become culturally. We can't even figure out that the use case for art is hanging on the wall for decoration. For a conversation piece. A few will have their name become known and make it into galleries.
Infinite supply means the value tends towards zero.Good luck monetizing anything with those economic characteristics.
30 seconds is probably a decent sweetspot imo.
If you are getting what you want most of the time, then you are a better 'prompt engineer' than I am.
For example it could take an old video game say morrowind and it could in real time patch the graphics onto the video screen. Or people could look at a video of themselves and it would update the style similar to a snapchat filter.
But the difference is academic; progress is so fast that it is reasonable to expect all these models will be obsolete in a year or two.
I'd love to read a less technical writeup explaining the challenges faced and why it took so long to figure out spelling. Scrolling through the paper is a bit overwhelming and it goes beyond my current understanding of the topic.
Does anyone know if it would be possible to eventually take older generated images with garbled up text + their prompt and have SD3 clean it up or fix the text issues?
I'm not sure about re-doing the text of old images - you could try img2img but coherence is an issue, more controlnets might help
We looked at different solutions extensively (https://medium.com/towards-data-science/editing-text-in-imag...) and ended up building a tool to solve the problem: https://www.producthunt.com/posts/textify-2
Eventually models will get to the point where they can do this well natively but for now the best we can do is a post-processing step.
I'm not an ML researcher but I can answer this. Note that this is not information from the paper, just my own findings from following the "scene".
We actually figured out spelling not long after diffusion models came out, the imagen paper that came several months before SD1 explained how they did it, which is the same technique SD3 and Dalle3 use, instead of using CLIP's text encoder, they use T5.
The reason why image models can't spell is the same reason why language models also have difficulty spelling. Tokenization. Simply speaking instead of seeing each letter individually, we split a sentence into sub-words, most commonly called tokens and that's what the model sees. The model never gets to see each letter individually. But it turns out that if you make the model big enough and feed it enough data, it actually learns how each token is spelled out.
Clip is both small (200-500M params) and trained on limited text data (only image captions). T5 is trained on a large corpus of data and is also huge (~5.5B params). This makes T5 the obvious choice if you care about spelling.
So why use CLIP in the first place? Simple: it's much easier to train on clip embeddings than T5 embeddings, not only are they smaller, but because clip is trained on text/image pairs and due to backpropagation, the text embeddings also contain a lot of visual semantic information. This simplifies a lot of the work the text->diffusion attention modules need to do. Another reason is that T5 is absolutely massive, it's 5x larger than the image part of the model and 10-20x larger than CLIP.
If you take a close look at diagram (a) in the paper you'll actually see that SD3 uses both CLIP and T5, not only that but they trained it in a way that makes the encoder used optional, so you can use the CLIP models only if you don't care about spelling and image composition (CLIP is also bad at understanding prompts), which is useful because most GPUs can't handle T5 on it's own. Tho I suspect someone will distil T5 so it becomes 5-10 times smaller than it currently is at a minimal loss on how good it is for prompting.
Why:
It looks generic enough to incorporated text encoding / timestep condition into the block in all the imaginable ways (rather than in limited ways in SDXL / SD v1, or Stable Cascade). I don't think there is much left to be done there other than to play with positional encoding (2D RoPE?).
Great job! Now let's just scale up the transformers and focus on quantization / optimizations to run this stack properly everywhere :)
YC started out with the intent to give young smart people a shot at starting a business. IMHO it has shifted significantly over the years to more what you say. We see ads now seeking a "founding engineer" for YC startups, but it used to be the founders were engineers.
I’d be happy if my government or EU or whatever offered cash grants for open research and open weights in AI space.
The problem is, everyone wants to be a billionaire over there and it’s getting crowded.
Bitcoin is trust-solved because of how the new blocks depends on previous blocks. With training data, there is no such verification (prompts/answers pairs do not depend at all on other prompt/answer pairs) (if there was, we wouldn't need to do the work of training the data in the first place).
You can rely on multiplying the work where gross variations are ignored (as you suggest): but that will take a lot more overhead in compute, and still is susceptible to bad actors (but much more resistant).
There is no solid/good solution - afaik - for distributed training of an AI (Open assistant I think is working on open training data?), if there is: I'll sign up.
Of course it gets harder as models get larger but distributed training doesn't seem totally infeasible. For example if we were to talk about MoE transformer models, perhaps separate slices of the model can be trained in an asynchronous manner and then combined with some retraining. You can have minimal regular communication about say, mean and variance for each layer and a new loss term dependent on these statistics to keep the "expertise" for each contributor distinct.
Do you want to 1. be right
or
2. stay in business
This is one of the reasons why OpenAI pivoted to be closed. Not bc of greedy value extractors; because it was the only way to survive.
Too bad they decided to get greedy :-(
I don't know you but I hope things work out :)
It's just hard being reminded that there's no escape hatch - we've welded them all shut for eternity. Being reduced to choices within a system but the choice horizon never extends to the system itself and won't within my lifetime makes me feel trapped.
Which is why they are not the future. A big model that can generate a picture about anything in response to any input makes for a great website. It generates lots of press. But it is not a reasonable tool for content generation. If you want to produce content in a specific area or genre, the best results come from a model trained or modified in the area. So the big generalized AI, if you use it, would only be the framework on which you built your specialized tool. Building that specialized tool, such as something dedicated to images of a particular politician, does not require huge amounts of computation. That sort of thing can and is being done by individuals.
I am waiting for a tool trained on publicly-accessible mugshots. It wouldn't be a very big project but could yield a tool to generate very believable mugshots of politicians.
Big generalist models are the future.
The big providers are all so terrified they'll produce a deepfake image of obama getting arrested or something, the models are so locked down they only seem capable of producing stock photos.
I think the content they are worried about is far darker than an attempt to embarrass a former president.
https://www.theguardian.com/technology/2023/oct/25/ai-create...
The internet has been taken over by Capitalist and have ruined the internet, in my opinion.
One has to wonder, why the delay?
Stability will release them in the coming weeks.
I wonder if anyone in Open AI openly says it "We're in for the money!"
The recent letter by SamA regarding Elon's trial had as much truth as Putin saying they are invading Ukraine for de-nazification.