Diffusion Bee: Stable Diffusion GUI App for M1 Mac
github.com
github.com
- No way to specify a seed. This is an important part of SD workflows, letting you redo an image with a slightly tweaked a prompt.
- No way to specify a custom model. Alternative models (such as Waifu Diffusion) are fun to play with too.
- No way to generate batches of images.
- No way to specify which sampler to use.
- No way to adjust the weight of specific sub-phrases, or use negative weights.
There's also no img2img yet, but it sounds like that's a planned feature.
Other GUIs such as https://github.com/sd-webui/stable-diffusion-webui, https://github.com/AUTOMATIC1111/stable-diffusion-webui have many more features - but not all will be trivial to port to M1.
P.S.: For the people asking - yes, it can do NSFW images. I checked.
There's https://www.charl-e.com/
I'm used to troubleshoot all kinds of stuff but I'm not a Python guy and wrangling with dependencies and virtual environments is not fun in my book. I've got lstein repo working but can't wait to have cleaner way of doing things.
Now it only generates black frames, even after being restarted. Not impressed at this point.
The options are sizes up to 768 x 768, "steps," and "guidance scale." You can do text-to-image or image-to-image.
Weird that it's considered a big challenge to fix… something in Metal? Core ML? SIMD libraries? Some translation layer? Something is not being said here.
which offers custom seed and batches of images (we are also working on parametric prompting https://github.com/breadthe/sd-buddy/discussions/12#discussi... )
i'd love to get to img2img and alt models next.
A couple of suggestions, maybe you can implement:
1) option to switch the conda environment since I had setup the dependencies in a specific environment (for me, the command is "conda activate ldm")
2) It would be nice to know what was the error when there is an error.
"Stable Diffusion Buddy is open source and free for personal use.
You may not use Stable Diffusion Buddy for any commercial purpose. That means you may not sell or profit in any way from the compiled app, from compiling the app yourself, from the source code, or a fork of it. The images generated by using this app do not fall under these limitations for obvious reasons."
$HOME/.diffusionbee
$HOME/Library/Application\ Support/DiffusionBeeAs a developer the best you can usually do is put some instructions somewhere (website, somewhere in a menu in UI), or do a package install that puts an uninstaller somewhere (more of the Windows way, but it breaks many expectations from users like having your app be one icon in `Applications`, and not it's own folder with your app and maybe a "Uninstall" app there too). For Aerial I give uninstall instructions on the website, but there's not much of a standard on how to handle this (that I know of, at least).
With the various security changes in "recent" macOS, you can't modify your bundle and download files inside it, as it would change its signature and break the notarisation system.
God I would hate to be whoever OpenAI blames this on.
I’m going to leave that DALL-E beta approval email unread like a food delivery recruiter email.
Now, I've seen an awful lot of programs written in Python that decide to force you to install them with virtualenv or pipx, but that's not Apple's doing.
Not true. There's no python2 anymore, but there is a python3. You can find it at /usr/bin/python3 (most likely your homebrew install preceeds it in your PATH)
This is/has been changing slowly, though. Things are much better then they were a year ago.
The best way to install Python on a mac is the downloadable installer from Python.org (!)
Features: - Full data privacy - nothing is sent to the cloud - Clean and easy to use UI - One click installer - No dependencies needed - Multiple image sizes - Optimized for M1/M2 Chips - Runs locally on your computer
Still, probably not too hard to build for iPad or iOS.
Another comment mentions RAM capabilities. Unfortunately that’s tied to the storage tiers instead of being something you can pick separately, so if you want 16 GB of RAM you have to buy the 1 TB or 2 TB models. Meaning for a 12.9” iPad Pro, if you want 16 GB you’re looking at an $1800 tablet. Not ideal.
M1 ultra was compared to RTX 3090 which was a larger stretch.
The M1 max deliver about 10.5 tflops The M1 ultra about 21 tflops.
The desktop RTX 3080 delivers about 30 tflops and RTX 3090 about 40.
Apple’s comparison graph showed the speed of the M1s vs. RTXs at increasing power levels, with the M1s being more efficient at the same watt levels (which is probably true). However, since the graph stopped before the RTX GPUs reached full potential, the graph was somewhat misleading.
The M1 max and Ultra have extra video processing modules that make them faster than the RTX GPUs at some video tasks though.
Seriously though, I imagine this is less a case of whether this specific implementation permits pornography, but whether any porn was included in the dataset it was trained on. No matter how good AI is, it only knows what it knows.
It mostly understands what naked people look like, but the images I've generated involve a lot of accidental body horror. You get a lot of people with extra arms, weird eyes, or body parts in the wrong places. The fact that they explicitly removed porn from the training set comes through pretty clearly in the model.
I suspect it could be improved a lot with some specialized retraining. As far as I know, nobody has done that work yet.
The official implementation has a second model that detects pornography, and replaces outputs including it with a picture of this dude https://www.youtube.com/watch?v=dQw4w9WgXcQ (Not kidding). Removing that is a really simple a one line change in the official script.
I was amused by this when reading the source. Here’s the function that loads the replacement.
https://github.com/CompVis/stable-diffusion/blob/69ae4b35e0a...
It looks like removing line 309 of the same file would disable the check, but I haven’t tried it.
But the second point here is also wrong: the whole reason these models are interesting is because they can generate things they haven't seen before - the corpus of knowledge represents some type of abstract understanding of how words relate to things that theoretically does encode a little bit of the mechanisms behind it.
For example, it theoretically should be able to reconstruct human like poses it has never seen before provided it has examples of what humans look like and something which transposes to an approximate value - an obvious example in the context of the original question would be building photorealistic versions of a sketched concept (since somewhere in it's model is an axis which traces from "artistic depiction of a human" to "photograph of a human" in terms of style content).
Of course, most people aren't very good at drawing realistic human poses - it's a learned skill. But the magic of deep learning is really that eventually it doesn't need to be - we would hopefully be able to train a model which can be easily copied and distributed which represents the skill, and SD is a big step in that direction (whether it's a local maxima remains to be seen - it's dramatic, but is it versatile?)
It's not good at penises or vaginas, just breasts and butts. I can find what I like pornographically without ai. But the ducking and dodging around nudity and sexuality is childish and tiresome and we ought to discuss this topic as disinterestedly and nonchalantly as we do the other things it excels or struggles at. Penises and vaginas are not somehow more vile than other things.
Sure enough, sexual content ought to be age-gated and there is potential for abuse (slapping a person's face onto explicit imagery without their consent isn't cool yall). But are we working toward AGI or not? Because at some point it's gonna have to know about the birds and the bees.
:) made me chuckle. AGI would need to know how to be evil, as well, right?
[1] https://www.allure.com/story/vagina-vulva-difference-planned...
we're creating the speakwrite machine.
No, it's essentially generating mashups of its training data, which can be very interesting.
So a model that hasn't been trained on a lot of porn will of course do a very bad job at generating porn.
It requires a good steer and prompting is a clumsy tool for fine tuning - it's adequate for initialization but we lack words for every shade of meaning, and phrase weighting is pretty clumsy too, because words have a blend of meaning.
The very fact that the model is interpolating between things in the latent space probably explains why its images haven't been explored by human artists before: because there is a disconnect between the latent space of the model and genuine "latent space" of human artistic endeavor, which is an interplay between the laws of physics and the aesthetic interests of humans. I think these models know very little about either of those things and thus generate some pretty interesting novelty.
Aesthetic choices like colour and shapes and composition combine with literal representations, facial emotions, symbolic meanings and so on. AI art so far feels quite shallow by this metric, usually only hitting a couple of notes. But sometimes it can play those couple of notes very sweetly.
Its like if Photoshop broke itself when you tried to modify or create anything nude or provocative, as the default
That would just be weird and thats what these AI software devs have done
so everyone patches that contrived feature flag, but nobody knows if they patched it
The filter should be easy to remove and there are already people who simply removed the filter.
https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/im...
This is just a small fraction of the imagery and it does include pornography.
I respect that if you're online and reading HN, you're probably mature enough to handle seeing pornography. So if you're curious to see some of the training data that made it in: choose from the dropdown "-column-" and change it to "punsafe" and set the value in the adjacent field to "1", then press Apply.
Obviously this will show pornography on your screen.
An article which talks about the imagery and how this browser came about is here: https://waxy.org/2022/08/exploring-12-million-of-the-images-...
If you are on twitter and would like to share : https://twitter.com/divamgupta/status/1569014206912929796
Edit:
Yep, that's it:
"you can't (yet)
this is reiterated here like a 100 times at this point
images and prompts were scraped from dicord bots
on the official discord
as there's no copyright on the images"
Not one click install by a long shot, but the documentation is pretty clear to follow. Anyone with a bit of CLI experience can do it and if you don't have that, this is a great way to kind of stumble your way towards something working and learn in the process...
https://github.com/lstein/stable-diffusion/tree/development/...
If people have any sort of feedback I'd love to hear it, or if people have some specific features that are missing from the other UIs :)
Hopefully I can do a Show HN in the future, when/if there is a free version for people to play around with, and get some really good feedback that way.
I generate one image in about ~3 seconds with the DDIM sampler, 20 steps, on a RTX 2080Ti (~8it/s). The video on the Patreon page is sped up as it's not very interesting to sit and watch renders haha.
Although, some of the users who started using my UI weren't using the fork my app connects to, and were surprised it was a bit faster than what they were using before, so maybe you can give it a try. The repository is https://github.com/lstein/stable-diffusion
It's true you have to use the code from git rather than a release, but that's not hard.
https://github.com/nlothian/m1_huggingface_diffusers_demo is my clean demo repo with a notebook showing the usage. The standard HuggingFace examples (eg for img2img[2]) port across with no trouble too.
[1] https://github.com/huggingface/diffusers
[2] https://github.com/huggingface/diffusers#image-to-image-text...
The fact that I can do this on commodity hardware on a 4GB model. A model that understands text and visual images, just absolutely blows my mind.
I almost feel like in a new future, a 100GB model may be able to offline handle speech -> text, video -> live scene graph. A robot that could base level physical understanding of our world like a 4 year old does. (objects, their relationship to other objects and behaviors)
System preferences "Software updates" tab says I'm on the latest (12.4) and there are no updates for me to install.
How am I supposed to try this?
What gives.
edit: I went to the App Store and found the listing for Monterey and clicked "GET" which opened the software update dialog with 12.5.1 and the option to upgrade.
According to others, it’s about a minute or less with 16GB.
https://nmkd.itch.io/t2i-gui https://github.com/n00mkrad/text2image-gui
We're at a turning point.
Since it's a Mac app, I have to wonder if it could stick the prompt, steps, and guidance into the notes field of Get Info? I find I'm generating a lot of relatively low guidance (I'd love a 6.5 option) images and iterating on the prompts with an eye to what it's suggesting to the algorithm. As such I have no way to closely track what prompt was active on any output as it changes so often.
I strongly suspect the real merit of this approach is not the crowd-pleasing, 'set very high guidance on some artistic trope so it's forced to fake something very impressive', but rather the ability to integrate a bunch of disparate guidances and occasionally hit on a striking image. It's like the harder you force it into a particular mold, the more derivative and stifled its output becomes, but if you let it free associate… I'll be experimenting. Seems like getting the occasional black image shows you're giving it the freest rein.
Looking forward to 'image to image' a lot. I assume the prompt still matters, as it's fundamental to the diffusion denoising? Image to image means iterating on visual 'seeds'.
I've seen talk of textual inversion training: it would interest me greatly to be able to generate objects and styles and train a personal version of SD in a sort of back-and-forth iteration. The link to language is really important here, but so is the ability to operate as an artist and generate drawings, aesthetics and so on, to train the model. I did 440 episodes of a hand-drawn webcomic once, which had recurring characters and an ink-wash grayscale style I gradually developed. That means I have my own dataset, which is my own property, and certainly didn't make it big enough to make it into Stable Diffusion like say Beeple did.
Interesting times for the cybernetic artist. Basically computer-assisted hallucinatory unconscious, plus computer-assisted rendering. You could feed all of Cerebus (Dave Sim and Gerhard) into a model like this, panel by panel, and you'd probably get a hell of a lot of Gerhard out because so much of the panel area is tone and texture from him…
On a M1 Ultra, it takes 12 seconds with 64GB RAM (same settings). While computing, the Mac Studio is pulling 100 watts of power.
I've been running webui [1] on M1 MacBook Air 16GB RAM: 512x512, 50 steps takes almost 300 seconds. I'm suspecting that it is running on CPU, because the script says "Max VRAM used for this generation: 0.00G" and Activity Monitor says that it's using lots of CPU % and no GPU % at all. When M1 users are running stable diffusion, does the Activity Monitor show the GPU usage correctly?
I should max out on the Mac gpus?
Speaking of which, what is the license on this? (the electron app)
If we can have "keyframes" with prompts that would output a PNG sequence or video, that would be awesome.
Thank you. :)
No idea if it supports AMD GPUs though.
Sorry, Apple is nowhere close yet.