How to generate realistic people in Stable Diffusion
stable-diffusion-art.com
stable-diffusion-art.com
Try the single word prompt 'woman' and see what you get...
Have larger diffusion models gotten to synthetic training dogfooding yet?
The irony is that once we get there, we can address biases in historical data. I.e. having a training set that matches reality vs images that were captured and available ~2020.
It looks like the stock photo cover for that mandatory course you hated.
Even adding keywords like “everyday” doesn’t help. And a fear it’s going to be worse in a few years when this stuff constitutes the majority of the input.
"90s, single use camera, documentary, of anoffice worker in an open plan office, realistic, amateur photo, blurry"
Results: https://imgur.com/a/GJLqYft
Corporate accounts payable, Nina speaking. Just a moooment. https://m.youtube.com/watch?v=4s5yHUpumkY
Surely someone has done a paired kinematics model to filter results by this point?
Not my field, but I figured 11 fingered people were just because it was computationally cheaper to have the ape on the other side of the keyboard hit refresh until happy.
Yes.
For example, take a look at this LoRA which is one of my favorites: https://civitai.com/models/259627/bad-quality-lora-or-sdxl
This, along with a proper model and when prompted properly, will give you photos of people who actually look like real people.
By contrast, a "normal" random person is very easy to generate, but very difficult to keep consistent across scenes.
Wouldn't these kinds of negative prompts and tweaking break down if I wanted to plug in more varied descriptions of people?
I find it interesting to plug in colorful descriptions of person's traits from a novel for example, or of people actually doing something.
Using "ugly", "disfigured" as negative prompt probably wouldn't work then...
For the pictures in the article, my first association is someone generating romance scam profile pictures, not art.
It’s currently the best open weights model for prompt adherence too: https://imgsys.org/
This doesn't specifically measure prompt following but how good it is overall.
Sure the motorbike handlebar is messed up, as is one of the hands. And the torso doesn't quite match up with where the thigh is. And the reflections on the arms, the face and the cleavage all look differently lit. And the hair isn't behaving like wet hair. And the face has an uncanny-valley airbrushed look about it.
But those are all easily overlooked. Chuck this into a Facebook news feed and I think 70% of the general public would believe it was a photograph.
There's examples that don't do this but they are harder to find (and prompt for)
While people have tried going from a base model to a fine tuned model based on explicit images, I wonder if there are people are attempting to go the other way round (train a base model on explicit photographs and other images not involving humans; then fine-tune away the explicit parts), which might lead to better results?
IMO finetuning is a waste of time unless you do cutting edge stuff.
Loras are just as powerful as a finetuned model and you can train one in minutes even on consumer hardware.
Do you have some more details on training a LoRA in minutes? Last I tried, it took several hours on an RTX 3090, but I am sure there have been improvements since then.
Tbh I mostly use loras others trained, there are hundreds around for all kinds of things
Is this process objectionable to you? If so, why?
It’s also possible I’m simply misunderstanding what you’re objecting to.
I first heard the "we need naked images to generate good clothed images" when SD3 came out and suddenly it's everywhere. But it just doesn't make any sense to me and as far as I know it wasn't explicitly practiced in previous popular models.
I like how even with all the "please don't make it porn" terms in the prompt, you can easily see (by choice of dresses, cleavage, pose, facial expressions etc) which models "want" to generate porn and are barely held back by the prompt.
When one asks prompts for which it hasn’t seen in training data, the results start to look less realistic.
Have even seen adult video logos in generated images.
I very much strongly suspect AI is not what we think.
Eking out something "interesting" is difficult, especially with limited time and low-end hardware. Interesting is highly subjective of course. I tend towards the more artistic / surrealist style, usually NSFW. Only nudes, no pornography.
I've been experimenting these last few months with interesting generating images, trying to make them "artistic" rather than photo-realistic, or the usual bland anime tributes.
I usually pick a "classical" artist which already has nudes in their repertoire, and try to blend their style with some photos I take myself, and with the style of other artists.
Most fall flat, some come close to what I consider acceptable, but still have major flaws. However, due to my time and hardware constraints they're good enough to post. I use fooocus which is kind of limiting, but after trying and failing to produce satisfactory results with Automatic, fooocus is just what I needed.
I can't really understand why more people don't do the same. Stable Diffusion was trained on a long and diverse list of artists, but most people seem to disregard that and focus only on anime or realistic photographs. The internet is inundated with those. I'm following some people on Mastodon who post more interesting stuff, but they usually tend to be all same-ish. I try to produce more diverse stuff, but most of the time it feels like going against the grain.
The women still tend to look like unrealistic supermodels. Sometimes this is what I want. Sometimes not, and it takes many tweaks to make them normal women, and usually I can't spare the time. Which is unfortunate.
If anyone's interested, I post the somewhat better experiments in:
https://mastodon.social/@TheNudeSurrealist
Warning: Most are NSFW. But are NSFW in the way Titian's Venus, say, is NSFW.
For others like me who aren't familiar with fooocus, there's a lot of ai-related sites with that name. I believe this is the one parent is referring to:
Edit: except hypernetworks, I don’t think they are still relevant for most users.
How come this technology appears to be exclusively used to generate fake pictures of unrealistically good-looking women? And to what end..?
With AI image generation, I can start with the broad brush strokes for the character and then use AI to generate an image based on those prompts which can then help me further define the character to the point that it already feels fleshed out and real before I have even started the game.
I'm pretty pleased with what Copilot (using DALLE-3) spit out for my newest character, a Gothic-themed forensic medical investigator: https://imgur.com/a/bXHeqAX
Of course I have a battery of dozens of techniques and addons to improve the 1.6 models
Also controlnet is so useful.
I also generate 512x512 with 1.6 and then upscale to 1024 with iterative upscaling - adds a ton of detail. Then I can easily upscale to 4K.
If I do iterative upscale to 4K right away it takes like 20h on my m1. But that adds even more details.
And there are negative embeddings which are great eg badhandsv2
So those 4 have been the most impactful for me
More likely to generate 7 armed nightmare fuel monsters than a bag of plutonium.
I don't know if it was me misconfiguring it, or if the images in post were really cherry-picked.
You need to simulate poor lighting, dirt, soul, realistic beauty etc. Perhaps even situations that give a reason for a photo to be taken other than I’m a basic heteronormative woman who is attractive.
Actually it is in no single image in that blog post.
If you have a trained eye that is.