The 'Compact' Version of Stable Diffusion 3 Is Generating Monstrous Human Bodies
xatakaon.com
xatakaon.com
Also, I found this quote interesting: "...Stability AI’s insistence on censoring adult content from SD3’s training data". Da Vinci figured this out hundreds of years ago, that to draw accurate pictures of humans, you need to understand the human body.
FWIW this is better for my purposes, with the older versions I recall trying to generate illustrations of female scientists for use in professional settings, and having to do a lot of tweaking to avoid ahem chest issues.
It seems obvious that midjourney is trained on copyrighted material though. I've seen the latest version generate straight-up "Tom Cruise in Top Gun 2" and similar.
Once video data is involved though it seems likely that will change. And I reckon a side effect is that a lot of the trickier details will improve. It'll be a great experiment to figure out whether the models also need lessons in anatomy or whether they figure it out through pure observation.
The current problem for them really isn't the details in isolation but rather cohesive details throughout the entire picture in one attempt. It's very lacking and requires a lot of manual input, filtering, reliance on multiple tools, etc.
But that should not be the case. A human body is not more complicated than a horse’s or a cat’s, and those are usually much better. There really is a problem in our relation with our own bodies.
>But that should not be the case.
What is your point here? Good AI should not make mistakes? Or good AI should not handle those details?
Go ahead and ask some human artist how difficult it is to draw hands.
Yeah, I should have been clearer. If you release a generator model for the general public, it needs to work. You cannot say that the user need to fix it themselves or learn what inpainting is. From a user’s point of view, it either works or is broken. That’s ok for specialist tools or technical audiences that love playing with the bleeding edge, but that’s about it.
Anyway, I hope standards will change in the US faster than the rest of the world finishes shifting to those utterly stupid standards.
I live in Berlin, and outside my apartment there is a spinning cube of adverts; the face of that cube advertising for Dildo King is between the face advertising for family cargo bikes and the face advertising for Edeka (a supermarket). There are also several nudist beaches within the city limits.
From outside the USA, I sometimes hear things such as Florida wanting to treat cross-dressing as inherently sexualised and therefore criminal if done in the presence of minors. (Was that a true story? When I search for it, I get transgender issues, rather than cross-dressing, so I can't find out if Hillary Clinton's trousers are a literal fashion crime in Florida).
it's interesting to consider the implications of Taleb's essay in light of the companies trying to make globally inoffensive LLMs.
Unstable diffusion is where to go if you want to get into the NSFW community.
These images look like someone spliced salamander regeneration DNA into a human, cut them into many pieces, and connected the wrong thing to the wrong thing all over the place. Maybe that's what the AI wants to do to us, AM style (I Have No Mouth and I Must Scream)?
Can finetuning fix this?
wonder what the right checkpoints are. I could have sworn sd 1.5 will sometimes generate 3 flawless arms.
Cur that noise out.
What SD3 has done here is "good art," as modernity defines the term: It's unique, it's technically adroit, and, most of all, it's shocking, thought provoking, unsettling. You can only wonder what went through the mind of the "artist" -- what it was trained on, how it attained such a strange result. This is at least gallery-tier if not worthy of MOMA, and I'm 100% serious about this.
So it kind of makes sense that they would go with a conservative model that was virtually immune to the AI porn / deepfake panic.