Easy Stable Diffusion XL in your device, offline
noiselith.com
noiselith.com
Pros:
- seems pretty self contained
- built in model installer works really well and helps you download anything from CivitAI (I installed https://civitai.com/models/183354/sdxl-ms-paint-portraits)
- image generation is high quality and stable
- shows intermediate steps during generation
Cons:
- downloads 6.94GB SDXL model file somewhere without asking or showing location/size. Just figured out you can find/modify the location in the settings.
- very slow on first generation as it loads the model, no record of how long generations take but I'd guess a couple minutes (m1 max macbook, 64GB)
- multiple user feedback modules (bottom left is very intrusive chat thing I'll never use + top right call for beta feedback)
- not open source like competitors
- runs 7 processes, idling at ~1GB RAM usage
- non-native UX on macOS, missing hotkeys you'd expect, help menu. electron app?
Overall 4/5 stars, would open again :)
If you feel this is unannounced self-promotion, yes, it is, and can be done better.
---
Also, for the "objective" comment, it meant to say "the original comment is still objective", not that you can be objective only by being a developer. Being a developer can obviously bias your opinion.
I was also surprised on how well it ran on my iPhone 11 before I replaced it with a 15 pro.
(Let me know if you're looking for some Product Design help/advice, totally happy to contribute pro bono. No worries if not of course!)
Unsolicited UX recs:
- strongly recommend a default model. The list you give is crazy long. It kind of recommends SD 1.5 in the UI text below the picker but has the last one selected by default. Many of them are called the same thing (ironically the name is "Generic" lol).
- have the panel on the left closed by default or show a simplified view that I can expand to an "advanced" view. Consider sorting the left panel controls by how often I would want to edit them (personally I'm not going to touch the model but it is the first thing).
You are doing great work but I wouldn't underestimate the value of simplifying the interface for a first-time user. It seems to have a ton of features but I don't know what I should actually be paying attention to / adjusting.
Is there a business model attached to this or do you have a hypothesis for what one might look like?
This could honestly be the excuse I need (want) to order an absolute beast of a macbook pro to replace my 2013 model.
The performance gains in recent models and PyTorch are currently outpacing hardware advances by a significant margin, and there are still large amounts of low-hanging fruit in this regard.
Who are the competitors?
InvokeAI: Apache license 2.0 (web-browser UI)
automatic1111: AGPL-3.0 license (web-browser UI)
ComfyUI:GPL-3.0 license (web-browser UI)
There's more, but I don't pay enough attention to it
https://noiselith.notion.site/License-61290d5ed7ab4c918402fd...
So yes, it is an electron app with svelte, headless-ui, tailwindcss etc
And if the defense here is "but Auto1111 and Comfy don't have as user-friendly a UI", that's also already covered. https://github.com/invoke-ai/InvokeAI
InvokeAI is installed via a script, sure, but it's also just a few clicks: download, extract, double-click on a specific file, enjoy.
comfyui: excellent for workflows and recalling the workflows, as they're saved into the resulting image metadata (i.e. sharing images, shares the image generation pipeline)
InvokeAI: Great UX and community, arguably were a bit behind in features as they were focused on making the UI work well. Now at the stage of bringing in the best features of competitors - Like you, I can easily recommend it above all other options.
Doesn't a1111 already do this? Theres a PNG Info tab where you can drag and drop a PNG and it will pull all the prompt, inverse prompt, model, etc. And then a button to send it to the main generation tab. It doesn't automatically load the model, but that may be intentional because of how long it takes to change loaded models.
Not that provides the same thing, no, largely because of fundamental design differences.
> Theres a PNG Info tab where you can drag and drop a PNG and it will pull all the prompt, inverse prompt, model, etc. And then a button to send it to the main generation tab.
A1111 by nature, has a bunch of disconnected operations in separate tabs and scripts. Even if the PNG captures all of a generation operation that would be executed by a single launch-button click, its not really equivalent to capturing a whole ComfyUI workflow, which can be the equivalent of a process which would be numerous different tasks in A1111 with manually shuttling data between tabs and scripts.
A1111 has a bunch of manual "send to X" buttons to do with the output of runs, so that they can be the input of another task, wherein in Comfy those operations are part of one workflow with a pipeline connecting the output of one to the input of another. And when saving generation data, those manual shuttle points in A1111 are barriers as to what is part of a single generation that can be saved.
I'm positive this can be done w/ Comfy too.
Yes, you can, and the workflow JSON format has a reduced "API form" that discards visual/UI related information.
Also, if you are using Python, you could do your automation in Comfy (as custom nodes) instead of outside, too.
I've personally been using SD.Next, which is a fork of A1111 with support for the diffuser backend, a cleaned-up UI, and also sometimes has support for newer things before A1111, though not always. It's plugin compatible with A1111.
There are a bajillion local SD pipelines, but this one is, by far, the one with the highest quality output out-of-the-box, with short prompts. Its remarkable.
And thats because it integrates a bajillion SDXL augmentations that other UIs do not implement or enable by default. I've been using stable diffusion since 1.5 came out, and even having followed the space extensively, setting up an equivalent pipeline in ComfyUI (much less diffusers) would be a pain. Its like a "greatest hits and best defaults" for SDXL.
The hundreds of python scripts and having the user to touch the terminal shows why something like Noiselith should exist for normal users rather than developers or programmers.
I would rather take a packaged solution that just works over a bunch of scripts requiring a terminal.
Look, DiffusionBee is still maintained but still no SDXL support.
Anyone who bet that the technology is done and it is time to focus on the UI is making the wrong bet.
git clone https://github.com/lllyasviel/Fooocus.git
cd Fooocus
pip3 install -r requirements_versions.txt
python3 entry_with_update.py
> pip3: command not found
Okay. I'll need to install it? What package might that be in, hmm. Moving on, I already know it's python.
> /usr not writeable
Guess I'll use sudo...
= = =
Obviously I know better than to do this, but very few people would. This is not 'dead simple'! It's only simple for Python programmers who are already familiar with the ecosystem.
Now, fortunately the actual documentation does say to use venv. That's still not 'dead simple'; you still need to understand the commands involved. There's definitely space for a prepackaged binary.
This said, it's nice when developers attempt to detect the executable they need and warn what package is missing.
Additionally, some package choices depend on hardware.
In the end, a lot of the more popular projects have "one click" scripts for auto installs, and there are some for Fooocus specifically, but the issue there is its not as visible as the main repo, and not necessarily something the dev wants to endorse.
SDXL in particular isn't one of those "compute light, bandwidth bound" models like llama (or Fooocus's own mini prompt expansion llm that in fact runs on the CPU).
There is a repo focused on CPU-only SD 1.5.
Can our entire field please realize that running this surveillance is a bad move, and just stop doing it.
If it isn’t explicitly surveillance, it could effectively be.
It does look bad that it bundles GTM, though, as a sibling commenter says.
Samples:
No, its not.
There are two text encoders, but they aren't really “prompt” and “style” inputs.
> and most other UIs dont implement the style prompting.
Most UIs default mode of operation sends the same input to both text encoders, but at least comfy has nodes that support sending separate text to them. OTOH, while there may be some cases where sending different text to the two encoders helps in a predictable way, AFAIK most of the testing people has done has shown that optimal prompt adherence usually comes from sending the same to both.
thanks for any tipp guys :)
I'd probably focus more on it being easy to install and use, as that's something that isn't done much. For me, if it doesn't have Controlnet, upscaling, some kind of face detailer, and preferably regional prompting, I'm out.
I also kind of wish all of these people that want to make their own SD generators would instead work on one of the open source ones that already exist.
While an app store might be a good idea, in a world with Auto111 and all of their extensions I think it's going to go over poorly with the Stable Diffusion community, for what it's worth.
I can see how something simpler might appeal to new users, even if it doesn't appeal to existing users.
It was weird when I was first playing with SD how many packages did severe phone home or vms or whatever instead of just downloading a bunch of stuff and running it.
I mean, really??
To me this is more like yelling "ROVER! COME HERE BOY!" at the top of your lungs.
OP is just offended by the image of an attractive woman, I guess. Apparently that's "creepy" now.
That is completely wrong to advertise it as “offline” if it requires an active internet connection to run.
On the first run it downloads about 30GB of data. I don't know if it would work offline on subsequent runs because for me it never ran again without crashing!
Also upon uninstallation it left behind all its data (not user data, mind you. But the executable itself, its python venv, its updater, and all the models. Uninstall basically just removed the shortcut in the start menu).
Then there's ComfyUI, the holy grail of complicated, but with that complication comes the ability to do so much. It is a node-based app that allows you to create custom workflows. Once your image is generated, you can pipe that "node" somewhere else and modify it, eg: upscale the image or do other things.
I'd like to see if Noiselith or some others offer support for SDXLTurbo -- it came out only a few days ago but in my opinion is a complete game-changer. It can generate 512x512 images in ~half a second on consumer GPUs. The images aren't crazy quality but that ability to make a prompt like "fox in the woods", see it instantly and then add "wearing a hat" and see it instantly generate again is so valuable. Prior to that, I'd wait 12 seconds for an image. Sounds like not a big deal, but the value of being able to iterate so quickly makes local image gen so much more fun.
When you want to apply advanced workflows and repetitive tasks or use something cutting edge then Comfy is handy.
Another A1111 alternative to try which focuses on prompt generation is:
If anyone has a spare thousand hours to kill, I would build that and connect it up with the various front-ends including ComfyUI, A111, etc.. not a small amount of effort, but it will be rewarding.
So, Civitai.com, if they had an API for the ob-site generation and training functions?
What is the catch?
IDK, this all seems weird considering there are four other really good projects that do all of these things already.
There are many great suggestions and links to other similar/better packages, so follow the comments for more info, thanks :-)
So now I have to contemplate shopping for a new mac TWICE in one year (never happened before).
An M3 Ultra might be a more reasonable comparison for the 4090.
Agree the prices are crazy right now, though.
One of my coller Q12023 ChatGPT experiences was having it help me "reason through" which machine was most "upgrade-proof," dollar-for-dollar.
Now in Q42023, I would definitely had made the decision to purchase an M2 Studio (base model) instead — those additional upgrades (VS M2Pro mini config, sim.) were much more cost-effective. Overall, I'm extremely satisfied with my M2Pro base model.
what is the best way to creat myself an "Modell" or "checkpoint" of my wished modell face?
i find tutorials but they seem complicated and outdated. thanks for any small tipp what things to research for actually method. Thanks
Literally the reason why I am coming to HN every day! Thanks devs :)
Apple Silicon M1, 32GB RAM, in any case.