Stable Diffusion launch announcement
stability.ai
stability.ai
I had a fair bit of fun with DALL-E, but it's very expensive and I found the product too much like a toy - a pair of plastic scissors made for children. Many times have I got my prompts blocked for apparent "scunthorpe" style filtering.
Also the creativity is a bit muted, and I'm convinced anything an American would describe as un-christian has been purged from its training data leaving an air of vapidity - I found myself wasting many prompts trying to get it to generate vomit for example.
And of course, the ironically named "OpenAI" being beaten to the punch for actually releasing something that isn't just an API.
So 9/11 is cool, but boobs apparently not. What a world we live in.
With DALL-E mini, the classic opening sentence of the dark tower "the man in black fled across the desert, and the gunslinger followed" produces evocative (if abstract) artwork. With DALL-E 2? You get a message that your prompt is inappropriate and your account will be reviewed for termination if you continue. So they don't want guns, okay whatever.
But it eventually impacts pretty much any concept you want to try. I was exploring scifi art and wanted to see what it would do if I tried to get a stylistic fusion of old-school soviet spacecraft aesthetics (exposed structures with bulbous pressure vessels housing the controls [among other things]) with western equivalents. Fusing disparate design philosophies in a way that actually feels creative is a task at which DALL-E 2 occasionally performs superbly and I was really looking forward to playing with the concept -- dreaming through the machine of a long-lost potential future.
Nope! Turns out "Soviet", "USSR", and IIRC now "Russia" are straight up banned words. I burned up quite a few prompts before figuring out that was the trigger.
And then there's the cost. During the closed beta they gave us what's now several hundred dollars a month in access. That translated to less than an hour a day of exploring DALL-E. It felt inadequate when it was free, and now they're asking for hundreds of dollars a month for it? No.
The pricing model makes sense if you're trying to generate bland corporate artwork for some random webpage but as a creative tool it just ensures that the only people really engaging with it have substantial financial backing for producing exclusively milquetoast ``artwork.''
Oooooh I can actually run this at home!
Someone else said GPU PyTorch on the M1 is far from ready so I'm wondering if this will be CPU instead.
I get in the region 2-3 minutes from Disco Diffusion etc on a mobile 3080.
From the repo:
"To prevent misuse and harm, we currently provide access to the checkpoints only for academic research purposes upon request. This is an experiment in safe and community-driven publication of a capable and general text-to-image model. We are working on a public release with a more permissive license that also incorporates ethical considerations."
Big kudos to GitHub, Papers-with-code, Huggingface, Google Colab, Replicate, Discord, probably many other tools, and then everyone playing a part (especially the disco diffusion crowd).
https://paperswithcode.com/ https://huggingface.co/spaces https://replicate.com/
Unrelated, but multimodal.art has been doing very cool work on building a whole little app you can run from a colab. But their models are pretty underwhelming at the moment.
[0] https://github.com/CompVis/stable-diffusion#image-modificati...
I think I heard the Stable Diffusion folks call DALL-E the McDonald's of AI art, and based on my experience, I agree.
There's a discussion of this in the VQGAN-CLIP paper, see in particular 6.1 "Efficiency as a Value" https://arxiv.org/abs/2204.08583
Disclaimer: I'm one of the authors of the VQGAN-CLIP paper and was tangentially involved with Stable Diffusion.