Stable Diffusion XL technical report [pdf]
github.com
github.com
• Earlier SD versions would often generate images where the head or feet of the subject was cropped out of frame. This was because random cropping was applied to its training data as data augmentation, so it learned to make images that looked randomly cropped — not ideal! To fix this issue, they still used random cropping during training, but also gave the crop coordinates to the model so that it would know it was training on a cropped image. Then, they set those crop coordinates to 0 at test-time, and the model keeps the subject centered! They also did a similar thing with the pixel dimensions of the image, so that the model can learn to operate at different "DPI" ranges.
• They're using a two-stage model instead of a single monolithic model. They have one model trained to get the image "most of the way there" and second model to take the output of the first and refine it, fixing textures and small details. Sort of mixture-of-experts-y. It makes sense that different "skillsets" would be required for the different stages of the denoising process, so it's reasonable to train separate models for each of the stages. Raises the question of whether unrolling the process further might yield more improvments — maybe a 3- or 4-stage model next?
• Maybe I missed it, but I don't see in the paper whether the images they're showing come from the raw model or the RLHF-tuned variant. SDXL has been available for DreamStudio users to play with since April, and Emad indicated that the reason for this was to collect tons of human preference data. He also said that when the full SDXL 1.0 release happens later this month, both RLHF'd and non-RLHF'd variants of the weights will be available for download. I look forward to seeing detailed comparisons between the two.
• They removed the lowest-resolution level of the U-Net — the 8x downsample block — which makes sense to me. I don't think there's really that much benefit from wasting flops on a tiny 4x4 or 8x8 latent tbh. Also thought it was interesting that they got rid of the cross-attention on the highest-resolution level of the U-Net.
It makes sense to have different criteria of what good is for each stage
"The structure of the work was usually as following: Rubens sketched and corrected the painting at the very end, and his "staff" was given the whole main stage. To increase the speed and efficiency of work, Rubens shared the duties: some of the pupils painted the background, others were focused on details — they worked on foliage or clothing, the master himself corrected the whole work and execute the most "important" parts — hands and faces."
Link: https://arthive.com/publications/2854~Rembrandt_the_teacher_...
Presumably the two parts of the SDXL model are complimentary: a first pass that's an expert on overall composition, and a second pass that's an expert on details.
Interesting - if it can be detected then it can be removed, right?
In theory they could also train the model to always add a watermark.
There are papers around describing how to bake a watermark into a model. Don’t know if anyone is doing that yet.
(part from the effort required to put it in place)
And yeah, you can see the watermarking code in the huggingface pipeline now. Its pretty simple, and theres really no reason to disable it.
If it's strictly a primitive watermark with absolutely nothing unique per mark, then it might be no big deal.
https://github.com/huggingface/diffusers/blob/2367c1b9fa3126...
I see no reason to disable it. Its fast, and its no more personal than basic metadata like image format and resolution... I can't think of any use case for disabling it other than passing AI generated images as authentic.
as for training the model to add a watermark, that's an interesting idea but I'm not sure it's feasible/not trivial to circumvent.
Quantization & subsampling do a hell of a job at getting "unwanted" information gone. If a human cannot perceive the watermark, then processes that aggressively approach this threshold of perception will potentially succeed at removing it.
Because our main interest is this meaningful content, it will be harder to scrub from the image.
It seems much more likely that it's their solution to detect and filter AI images from being used in their training corpus - kind of a latent "robots.txt".
A robust rotationally invariant repeating-but-not-obviously micro watermark would be quite a neat trick.
> the model may encounter challenges when synthesizing intricate structures, such as human hands
I think there's two main reasons for poor hands/text
- Humans care about certain areas of the image more than others, giving high saliency to faces, hands, body shape etc and lower saliency to backgrounds and textures. Due to the way the unet is trained it cares about all areas of the image equally. This means model capacity per area is uniform, leading to capacity problems for objects with a large number of configurations that humans care more about.
- The sampling procedure implicitly assumes a uniform amount of variance over the entire image. Text glyphs basically never change, which means we should basically have infinite CFG in the parts of the image that contain text.
I'm not sure if there's any point in working on this though, since both can be fixed by simply making a bigger model.
This is really cool.
So is the difference between the first and second stage. The output of the first stage looks good enough for generating "drafts" when can then be picked by the user before being finished by the second stage.
The improved cropping data augmentation is also a big deal. Bad framing is a constant issue with SD 1.5.
Supposedly the Clipdrop API was also going to support it. Clipdrop has a web page for SDXL 0.9 but nothing in the API docs about it.
Anyone know what the story is with the Clipdrop API and SDXL 0.9?
Using the same techniques, yes, this will fit in 8.
It should be easy(TM) with bitsandbytes, or ML compiler frameworks.
Btw very cool