Try Stable Diffusion's Img2Img Mode
huggingface.co
huggingface.co
https://github.com/hlky/stable-diffusion
It supports both txt2img and img2img. (Not affiliated.)
Edit: Incidentally, I tried running it on a CPU. It is possible, but it took 3 minutes instead of 10 seconds to produce an image. It also required me to hack up the script in a really gross way. Perhaps there is a script somewhere that properly supports this.
It's often easier to actually get models to run on CPU, due to simpler install configs and more available memory. Just painful to get a result out of it. Which might help keep the install simple, because it's not even worth optimizing
Not really.
https://news.ycombinator.com/item?id=32635086
I’m not even sure it works well - if at all - with 4Gb.
In any case, it’s impressive even if it takes minutes. And it’s not like you need to be there to make it work. You can create a list of prompts, let it do its thing and check the results later.
I’ve not tried that size but I tried 256x256 and it was too small to get interesting results - maybe there are some parameters that can be adjusted to improve it though.
I'm planning to try it out this weekend. But not really hopeful. I only have 8GB ram, and mine is a cheap Intel CPU that doesn't even have AVX.
GP talks about generating a single image while you talk about generating 4.
This is for the initial "exploration" step of the process. Once I like an image I typically play with the settings, then in the final step run with a large number of steps (and maybe even use upscaling).
So, default 512x512 size, 16 steps and default for the rest of the settings (I believe 7.5 scale, 0.75 strength).
Having said that, I also tried the official Docker image for stable diffusion and with the default values it generated an image in about 40 seconds.
I do runs at 384px by 384px, with batch size of 1. Sampling method has almost no impact on memory. Using k_euler with 30 steps renders an image in 10 to 20 seconds. The biggest thing that affect rending speed is the steps and the resolution, so 512x512 with C 50 using ddim is much slower than 256x256 with C 25 using k_euler.
The sampling methods run mostly in the same timelines, but the k_euler one can produce viable output at lower C values, meaning it is faster than the rest.
Don't add gfpgan in the same pipeline, as it takes more vram.
I'm running it on Windows 10 with latest drivers. I set the python process to Realtime priority in task manager (makes a slight difference!). Have not tried it on Linux.
I'm thinking about getting a 3090 so that I can make higher resolution images.
Gfpgan runs much faster for me 5 seconds per picture
edit: on second thought, they most likely rendered at 512px but then ran it through an upscaler model. I've been meaning to hook mine up but kinda forgot to try.
https://old.reddit.com/r/StableDiffusion/comments/wy7oa5/img...
https://old.reddit.com/r/StableDiffusion/comments/wyq04v/usi...
https://old.reddit.com/r/StableDiffusion/comments/wzlmty/its...
You can find the announcement tweet here: https://twitter.com/mishig25/status/1563226161924407298?s=20...
I ended up paying the $10 for Google Colab Pro and that's how I've been using this. Maybe I'll figure out how to get this working on my old 1080 TI to see if it's faster.
Anyway, for the one that I'm using which has a web UI, you can use this Colab link. It's pretty great! https://colab.research.google.com/drive/1KeNq05lji7p-WDS2BL-...
What I really wish was that the img2img tool could be used to take a text2img output and then "refine" it further. As it is, the img2img tool doesn't seem particularly great.
People on Reddit are talking about "I just generate 100 images and pick the best one"... but this is incredibly slow on the P100 GPU that Google has me on. Does this just require a monster GPU like a 3080/3090 in order to get any decent results?
Also how slow is your p100? I'm usually getting around 3 it/s. Maybe it's just because I'm used to disco diffusion where a single image took over an hour, but this is ungodly fast to me
https://news.ycombinator.com/item?id=32634139
complaining about the price of getting their images done. So part of it may actually exchanging money for a service. I bet a good chunk of it is investor cash though.
Their business model is basically supporting enterprise and private use-cases. For example, getting expert support for using these libraries, or hosting models and datasets privately. You can see more information about the pricing here: https://huggingface.co/pricing
They reached a $2 billion valuation after a recent round of funding so overall they're probably pretty flush with cash lol
In this case I had the prompt "cow chewing bone" with 4 squares representing the two pair of feet, the body and the head. None cared about chewing on a bone.
With DALL·E 2 I tried to get an image of a little girl building sandcastles and a monster threatening her:
"little scared girl building a sandcastle and a big angry monster is looking at her."
"little scared girl building a sandcastle six damaged sandcastles are to her side. a big angry monster is threatening her. it is dark." https://imgur.com/a/f5FFKOi
"little scared girl building a sandcastle with six damaged sandcastles to her side and a big angry monster threatening her"
Is there some kind of structure the sentences should follow?
Dalle is is bad at positional prompts. Ask for somethi g to be in the top rightbhand corner and it will appear bottom centre
Also, most of the good ones you see online are cherry picked from hundreds of runs, so set your batch size too 1000 and go to bed! After that, people then tend to run some of the good results through img2img, also with a lot of variations produced from a single image. Finally, some people also run them at higher resolutions if they have enough VRAM, as smaller resolution can distort or generate rubbish. For the messed up faces, they run it through gfpgan a few times to get prettier faces. Other than that, it is pure luck (using random seeds) to figure out what works and what doesn't. Use the 2 sites above to help you improve your prompts.
(meant in the context of stable diffusion)
That’s not the “incremental diffusion preview” it’s the “waiting in a queue” preview.
There is a fair amount of 3d models out there so it should be possible.
I suspect we will end up using multiple 2d images at different angles to generate a 3d model. I have seen this done before
there are a number of forks and auxilliary repos available that add UI, reduce memory requirements, etc.
Here are DALL-E 2 interpolations: https://twitter.com/model_mechanic/status/151297688118364569...
https://github.com/schmidtdominik/stablediffusion-interpolat...
Try again after seeing this post made into front page, it's actually working (and relatively very fast), just the loading screen is misleading ...
Instead we get Deep Learning AI's being used for generating and faking sentences, images, videos, voices, code and digital art all being trained on mountains of data in data centers all significantly contributing to the already burning up of the planet to no benefit and no efficient alternatives.
A dystopian future with deepfakes and easy fake news creation thanks to the technologists who helped create those things for others to create a new griftopia on top of that.
So you will own nothing, believe everything that you see on the internet, and be very happy.
Also there is pretty dystopian angle because these tools balantly stole all the work of artists by “learning” on their work (without pemission). Calling it learning is too nice we all know its just huge visual pattern copy mashup paste machine. People are not gonna invent with this - they just add name of artist they like to the prompt and get to call that their work.
Its not liberation - its exploitation. Its gonna destroy peoples lives by making their hard earned skills obsolete.
Then again it probably wont be so bad. Artists wont disappear and people wont suddenly become artists just because they can write into prompt.
But there is important difference between someone stealing and some algorithm copy replicating anything from the past in instant. One is a remix that brings something new (even if author doesn't want) the other in static its conservation. It will create side effects that will impact our (visual) culture. But who knows what those will be.
None of that mean that i deserve the hate. It's just a different opinion.
If we had time travel, they’d say now someone can kill our granddads and rewrite history. If we had flying cars, they’d say now our flats are worthless because everyone can fly by and gaze. If we had universal basic income, they’d say something about it too.
It’s really nice that we only got AIs that draws pictures and bitcoin that helps with stolen money.
My point being, there's enough good things and bad things to fit any mood and worldview. Everyone can basically pick as they'd like.
[0] something like this https://www.reddit.com/r/Futurology/
github with gui: https://github.com/hlky/stable-diffusion
dev repo (more features, may have bugs): https://github.com/hlky/stable-diffusion-webui
repo with docker: https://github.com/AbdBarho/stable-diffusion-webui-docker
colab repo (new): https://github.com/altryne/sd-webui-colab
can also run it in colab (includes img2img): https://colab.research.google.com/drive/1NfgqublyT_MWtR5Csmr...
demo made with gradio: https://github.com/gradio-app/gradio
However, "lot of child dying in a massive fire" is an a-okay query with some "interesting" results.
Prudes are such a weird bunch.
Of course that’s not what hentai means either.
First, the filters are only on the interface. The research is public. That's why there has been an explosion of new implementations. You are welcome to run the code yourself and make the horniest model you like.
But more importantly, this comes up every time, and everyone acts SO confused. But why? Artificial Intelligence is MASSIVE mainstream news. New advancements are arriving daily.
Do you want the news stories to be dominated by whatever heinous shit some bad-faith giggly teenager with a call of the void to offend as much as possible (i know, i was one) is able to generate?
Do you want the discourse to be dominated with "Won't someone think of the children?" blocking legitimate research and progress?
Europeans always like to dunk on Americans for being "prudes", but only a tiny section of Western Europe is progressive enough to not mind random nipples on their bus ads and television. All of Eastern Europe, Africa, India, and China are culturally still fairly conservative about sex - at least out in the open.
People think that preventing racist or sexual use of these models is indulging the prudish mores of America. But that itself is a very narrow perspective that ignores the perspectives of billions of people in the world.
I agree with you that this probably is required while the tech is new and then eventually won't be when everyone and their dog can run these things locally on their phone.
How is that the only alternative?
That discussion is over. You shouldn't care about a company trying to filter out bad shit on their own platform.
Pandoras box has already been open.
Your argument is entirely descriptive. We know why they do this. What should be argued here is how stupid and sad it is to block progress due to nothing more than religious or moral values.
I don't mean this in a moralistic sense. I have no qualms about images of blowjobs. I have no doubt that the porn industry has already deployed these models and is experimenting with this without filters. As a monetizable way to reduce the human costs, it's an entirely logical step, and adult performers should be as nervous as artists, and talking to lawyers.
But PROGRESS???
Moreover, sexuality is perhaps the most important theme in art, historically. It is a very legitimate thing to want to include in a tool such as this one. Its censorship is akin to "moral codes" of the past, completely regressive.
While I am entirely pro liberty in terms of consumption of pornography, it is not a replacement for human interaction, and can be used in conjunction with other tools to dehumanize and isolate individuals. Some individuals that then make negative contributions to the rest of society at large. These are not isolated incidents, and they've been rising: https://en.wikipedia.org/wiki/Misogynist_terrorism
There are many kinds of porn - some that exploit women, some that empower them in their production. And there are also many kinds of products - ones that focus on a human connection (whether vanilla or kinky) and ones that don't.
I would absolutely worry that artificial porn would be good enough to meet some of these demands but not most, and would ultimately be a net negative for society.
> Moreover, sexuality is perhaps the most important theme in art, historically.
I mean, it's demonstrably not, religious art imagery dominates by volume. But that's also not necessarily important - that's just who happened to be patrons of arts and had the means to commission them.
I'm not gonna deny that humans are horny and want to make lots of sexy art.
But you're also not going to get porn on basic cable. You have to seek that out extra for yourself, and that's going to be the case with AI-generated porn art also.