Charl-e: “Stable Diffusion on your Mac in 1 click”
charl-e.com
charl-e.com
This is THE thread to bookmark for many more resources: https://np.reddit.com/r/StableDiffusion/comments/xcrm4d/usef...
Most Python projects install with a matter of a single pip command.
Somewhere along the way, they became so complex as to require special installation programs and the like.
I'm not very familiar with AI and SD in particular, but from what I understand, this stuff is mostly-pure maths, so it shouldn't be a difficult thing to package and make portable. I know the models are rather large, but that's not really any additional complexity.
Software in general? Yes. But of you tried to leave the path the manufacture prepared, then you entered a world of pain. I remember how difficult it was to connect my first smartphone around 20 years ago to my Windows PC using ActiveSync to achieve a synced calendar. Just one example of: there was no download and run solution for processes that seem simple today.
This is actually a solved problem. But it's been solved lots of different times in different incompatible ways which tend to clash on any individual computer.
It's not the case that useful software was ever self contained. If you recall trying to do anything online in the early to mid nineties, you'll remember how complicated it was to use almost any website and how much manual fiddling and configuration was involved to get online.
Ignoring the internet, early games and graphical applications were a mess of settings and configuration. Even today you often have to tune graphics settings to get playable game performance on anything but top line hardware.
All respect due to them building something incredibly advanced, but you should view this as the place where the science has done its part and the software engineering is just getting started.
Why is there precisely one test (that has nothing to do with the core functionality)? Why is the Git history full of things like “finish”, “correct merg”, “fix more”, “add code”? Where is the linting config? Why are there print statements everywhere? Why does non-UI code have UI code embedded in it? Why is there random code commented out? Why is there no consistency across the codebase? Why is everything written as if they’ve never seen a Python project before (comments that should be docstrings, docstrings that should be comments, print("WARNING: ") instead of using logging or warnings, underscores in CLI flags, no shebangs in scripts intended to be run from the command line…)
Not all Python software gets all of this right, but it’s incredibly rare to have so many misses, even for hobbyist developers. Unless it’s code released by researchers, etc. It’s pretty typical in that context.
> This is not a project made by people without dev skills.
No, but it is a project made by people lacking software engineering skills, which is a distinction I drew in my earlier comment. Like I said, they can write clever code, but there’s a difference between bashing on code until it works and building it properly. This is the kind of codebase you get when you have people who have been writing code for a long time, but never in the context of a software engineering project alongside experienced software engineers they can learn from. Put them on a team like that and they’ll be forced to unlearn most of these bad habits fast because they’d never get their pull requests approved otherwise.
I’m trying not to be harsh – I understand that this code is more of a code dump from researchers than a real software project, and they’ve done some incredibly clever things here – but if somebody is suggesting that Python projects in general are like this, it really should be pointed out that this is not in this slightest bit representative of a typical Python project.
Stable diffusion already "happened" for Krita
I can't believe people endure this stuff on a day to day basis, I dread it every time, the fact that different versions of packages can't co-exist and like installing something can downgrade my setuptools which then breaks my whole installation. Not even wrapping this all up within conda solves this stuff it just means you can burn the whole thing and start over easily.
Maybe it's user error, but I never encounter these problems anywhere else.
1) python appeals to a lot of people that work in development-adjacent industries (like AI). These people don’t usually have to care about packaging
2) Python has gone through many outdated forms of packaging
3) The zen of python seems to have encouraged everyone to install third party libraries for the smallest of tasks (implementing retries, formatting phone numbers, etc). These small packages often have only a few number of maintainers who end up dropping off the map.
Modern python package management works pretty well, but there’s so much debt in the ecosystem I’m not sure when it’ll be better for end users.
[1] https://github.com/lstein/stable-diffusion/blob/main/docs/in...
I've been running the development branch, which has been working fine for me, but I've also rebuilt the dependencies a few times just to be certain.
One protip is to use symlinks for the training files.
For A.I. stuff I actually don't judge, these scripts are written by people who specialise in other things than software engineering and they simply put together some code to run their algorithms and as a result they are poorly engineered in many aspects.
https://github.com/divamgupta/diffusionbee-stable-diffusion-...
Today, these models are far ahead of the trademark attorneys, but there are powerful interests that are going to want to litigate the inclusion of these entities in the trained models themselves.
So the art copyright angle should not be the only one taken into consideration.
I'd recommend keeping a prompt list and finding what does/doesn't work for what you're after. Try shuffling the order of your prompt - the order of the tokens does matter! Repeat a token twice, thrice, hell make a prompt with nothing but the same token repeated 8 times. Play around with it! If you find an image that's very close to what you want - start generating variations of it. Make 20 different variations. Make variations of the variations you like best.
Also the seed is very important! If you find a seed that generated a style you really liked take note of it. That seed will likely generate more things in a similar style for similar enough prompts.
It's a semi-creative process and definitely takes some time investment if you want great results. Sometimes you strike gold and get lucky on your first generation - but that's rare.
https://www.reddit.com/r/StableDiffusion/?f=flair_name%3A%22...
Also, this post refers to a large number of relevant tools to use as well:
https://www.reddit.com/r/StableDiffusion/comments/xcrm4d/use...
https://dallery.gallery/the-dalle-2-prompt-book/
If one prompt doesn't work, try writing it in another way. Sometimes it helps to write things in multiple ways in the same prompt.
What kind of computer specs would be required to generate typical SD images in less than a second?
Probably overkill and could get away with something like a 3060 or so, but the 24 GB of VRAM come in handy if you want to generate larger images. I pushed it as high as 17 GB on some recent runs.
It's worth noting that I'm on a 5800X as well, I'm sure.
What's the advantage of using img2img as opposed to iterating on the seed value?
I guess I'm just manipulating probabilities in my favor?
Congrats to the stable diffusion team for their openness and inclusiveness!
Is there a trick to it?
There's also a knack for writing the prompts, generally you want to write your prompt as a list of short sentences. Don't make your prompt too short. Use concrete and clear concepts, beginning with your main subject, and describe their relation. You can also qualify the background, the mood, the material and the style among other things. Generate more than one sample - I normally go for five or ten samples so there are better odds of getting one that works, since the AI could try to interpret the prompt in different ways that make sense (but not to a human).
I can't try it right now, but you might try something like: "A man is standing on the street, close up. The man has an iphone in his hand. The man is Mark Twain. Regular city background. Impressionist. Detailed."
Tweak a few times and more often than not you'll end up with something satisfying.
Like what were you hoping for? and add terms that will drive towards that.
My comment was pretty clear, not sure why the two responses I got decided to give me advice about prompts.
SD is much more raw.
(and actually made the code independent of cuda)
EDIT: I take it back - all the menus are the generic Electron ones, so it is quite possible that the author is finding this part tricky.
Is it... doing anything? Do I just need to wait 10 minutes? 20?
Stable Difussion isn't so heavy... mostly you are limited by how many steps you want to do.
An image without “crisp” pixels that create lines is acceptable. A basketball with an extra/missing blue pixel is acceptable.
With words, they either haven’t invented Diffusion yet, or the nature of the problem is too hard for Diffusion.
With text you can’t have the model Diffuse to “teli ne abouf a fdog?” “The dod is graen and has stix leg5”. It’s just too obviously wrong.
Meanwhile, the “fuzziness” allows Diffusion models to be smaller, compared with models that need precision.
> Will this be available on Intel Macs?
> Yep, I'm working on making it compatible with older Macs.