Stability AI Launches Stable Diffusion XL 0.9
stability.ai
stability.ai
I think Stability is ostensibly showing that the images are closer to the prompt (and the left wolf in particular has some distortion around the eyes).
Comps here - https://imgur.com/a/FfECIMP
Does anyone have comparisons of how the model does on specific artist styles?
Simple prompts like "By $ARTISTNAME" worked very well in SD v1.5, and less so in v2.x, depending on the artist in question.
I'd say the blur effects on the left images are much cleaner as well. There are some weird artifacts at the fringes of objects in the earlier version.
I think the coffee cup looks better in the right phot, it seems a tad bit more real to me.
Like you I much prefer the alien photo on the left, but the photos are so stylistically different I'm not sure that says anything about the releases' respective capabilities.
It is hard to describe but there is a very unnatural "sheen" to the images on the left.
The SDXL 0.9 images look more photo realistic but they still aren't quite at the level that midjourney can do.
The best example is the wolf's hair between the ears in the SDXL 0.9 image. It is just a little too noisy and wavey compared to how a real wolf photo would look. Midjourney 5.1 --style raw would still handily beat this image if making a photo realistic wolf.
The jacket on the Alien in the SDXL 0.9 image also has too much of that AI sheen but it kind of works in this image as an effect for the jacket material so not really the best example.
The coffee cup isn't very good on either of them IMO. The trees on the right are still not blurred quite right. They are hiding the hand with this image on the right too. You can see how bad the little and ring finger is on the left image.
Obviously, this is all very nit picky.
The real comparison should be with SD 1.5/2.1, and is WAY better.
In the first example, the second image is more representative of Las Vegas for the foreigner I am, but none of them hav ethe scratchy found film requirement
In the second example, both fit the prompt, but the first image look more coming from a documentary than the second one
in the third example, the hand from the second picture looks much better
Also, they’ve (re-) established a universal law of AI: fuck it, just ensemble it
Not sure if true, sounds plausible tho.
The Stability AI API/DreamStudio API is slightly different. Yes, it's confusing.
I read this as: commercial use through our API now, self hosted commercial use in July.
Im just looking forward to the custom LoRA files we can use with it :D
RIP my 1080 TI.
Does anyone know what specific feature they need which 20+ cards have and older ones don't?
Edit - it’s not the RAM. 1080TI has 11GB and this press release says it requires 8. So I’m going to speculate that it’s because 1080 lacks tensor cores compared to the 20x’s Turing architecture
An 8GB requirement kinda sounds like they have already quantized the model though.
My guess is AMD users will eventually get low VRAM compatibility through Vulkan ports (like SHARK/Torch MLIR or Apache TVM).
Then again, the existing Vulkan ports were kinda obscure and unused with SD 1.5/2.1
We had this for SD 1.5, but it always stayed obscure and unpopular for some reason... I hope its different this time around.
Additionally, for Apple Silicon you likely need 64 GB RAM (since CPU/GPU memory is shared) which is expensive.
Despite its powerful output and advanced model architecture, SDXL 0.9 is able to be run on a modern consumer GPU, needing only a Windows 10 or 11, or Linux operating system, with 16GB RAM, an Nvidia GeForce RTX 20 graphics card (equivalent or higher standard) equipped with a minimum of 8GB of VRAM. Linux users are also able to use a compatible AMD card with 16GB VRAM.
I’m guessing that it will work eventually, though I’m not sure who will make that happen.
If they don't so it for SDXL, the port will probably take awhile (if it happens at all).
This is critical, for legality of use, ethics concerns, and the quality of the output (as overly zealous filtering can degrade the model like it did for SD 2.0).
AFAIK only SD can be run locally?
I do expect there are other bases out there, but haven't seen any of quality yet.
Before this release (XL 0.9) it's been unclear how much of the SD quality was in-house or came from their prior collab with Runway/Heidelberg.
(However, I thought Midjourney definitely was at some point)
I don't like how SD consolidated around the A1111 repo. The features are great, and it was fantastic when SD was brand new... but the performance and compatibility is awful, the setup is tricky, and it sucked all the oxygen out of the room that other SD UIs needed to flourish.
If you want another UI to flourish, clone both it and A111, copy and paste the bits from A111 you’d like to have in yours (with attribution) and push it up along with any features you personally want.
That does require developer time, and developers may converge on a popular implementation with good tests and lots of features as it’s easier to contribute.
The bottleneck isn’t really the community though, it’s the developers.
I worked trying to add torch.compile support to A1111 for a bit, fixing some graph breaks locally, but... It was too much. Some other things, like ML compilation backends, are also basically impossible.
Facebook's AITemplate backend even supports long prompts now.
- The codebase is cleaner more hackable, and (compared to base SAI code) more performant.
- HF continues to put lots of work into optimization and cleanup. For instance, they ensure there are no graph breaks for torch.compile, and work with other hardware vendors for thier own SD implementations.
I ended up using https://github.com/easydiffusion/easydiffusion
which has served me well so far.
While it has a fraction of the features found in stable-diffusion-webui, it has the best out of the box UI I've tried so far.The way it enqueues tasks and renders the generated images beats anything I've seen in the various UIs I've played with.
I also like that you can easily write plugins in Javascript, both for the UI and for server-side tweaks.
I use A1111 as a tool, but if I want to goof off, I queue up a bunch of prompts in Easy Diffusion and end up with a gallery built in real time. Its smaller feature set make it great for that.
It would basically be a rewrite, if I were to guess... And at that point they mind as well port everything to diffusers.
Can I do dreambooth here? If so, what commands do I use?
Dreambooth is gonna require an A100 I think... I doubt it will work on the free (16GB VRAM) Colab instances.
It's also a foundational model, not a finished product, and MJ will possibly use it, like they did in v4 with SD 1.5.