Stable Diffusion 2.0 and the Importance of Negative Prompts for Good Results
minimaxir.com
minimaxir.com
https://twitter.com/emostaque/status/1596864150134984705
> Current -ve prompts: ugly, tiling, poorly drawn hands, poorly drawn feet, poorly drawn face, out of frame, mutation, mutated, extra limbs, extra legs, extra arms, disfigured, deformed, cross-eye, body out of frame, blurry, bad art, bad anatomy, blurred, text, watermark, grainy
If you want a really good picture of an imaginary person it helps if you use "extra limbs, extra legs, extra arms" as negative prompts!
// The following code does not contain any bugs:
<tab>
// this fixes [the bug in the previous code]
and it worked. :)// This method will determine if input program halts or loops
I’ll reply here once I have the code generated.
/* drunk, fix later */
Might try: // lgtm!
All in all I hate it because the prompts I see are things like "cyberpunk forest by Salvador Dali". You've got a tool that gives you the power of Gandalf and you prompt that?
Which works, because most people on the internet tend to be detail-oblivious!
https://usercontent.irccloud-cdn.com/file/2csfvKjL/image.png
Which is about what you'd expect from a generator that understands patterns, but not meanings.
They miss on things that cannot be, because they don't understand things or rules, only patterns.
If you mean locally as in the size of a hand being right while holistically the person is wrong, no.
The overall images "tend" to be right (once you grasp prompting) and elements, even appear right at first, but if you focus attention on those elements, they are often not quite right.
So perhaps it's the definition of local and holistic.
Looking at the woman on the boat [0] closely, I would still 100 % believe that’s simply a still from a movie, probably from the 90s.
[0]: https://nitter.kavin.rocks/pic/orig/media%2FFihVvliXwAEdw0y....
That's one of the better prompts I've seen. Dissimilar but really strong aesthetic styles a skilled human could mesh pretty well, interesting images and shows up the strengths (some of the forests are really good, and the ones without trees are pleasantly foresty nevertheless) and weaknesses (it fails completely on 'cyberpunk' and 'Dali' once you start adding other parameters that influence the visual style) of the model.
Plus I'd be much more likely to end up with a calendar of "cyberpunk forest by Salvador Dali" images on my wall than "Mickey Mouse in a tuxedo with a cigar"
I'm pretty sure this problem is not hard to fix in the long run, though.
I was able to come up with someone paddling a canoe in a Turner seascape. The only thing I couldn't get right was a proper canoe paddle and paddling motion but everything else was pretty much perfect.
Specific common keywords like "amputated" may have a positive impact, though. Hard to tell. Doing apples-to-apples comparisons with negative keywords is challenging because even a single extra keyword tends to completely change the image.
One thing that SD really impressed me by, though, is its understanding of symmetry. "Symmetrical composition" is an incredibly powerful phrase: https://imgur.com/a/lioJ8ak
And it does, indeed, extend to anatomy as well – "symmetrical eyes" can help a lot, while "symmetrical arms" renders people with their arms raised or outstretched.
I decided to add a negative prompt. With a bit of experimentation I realised all the "bad" had no effect. However, "blob" actually made most of the deformities go away and "amputee" did help against partial limbs being generated.
Something that worked even better was replacing "gymnast" with "athletic man"/"athletic woman" in the positive prompt.
Take the negative prompt "bad hands". The AI doesn't know what bad hands are, that's a human concept. But it does know what hands are, so it hides them. In the example image the hands, arms, and feet are all hidden.
In theory, using the negative prompt "hands" would be just as effective.
I'm not an expert, but I was given the above explanation by someone who knows a lot more than me and it makes sense.
1: SFW, personally can't agree with bad_hands tag https://danbooru.donmai.us/posts/5797703
I imagine that on the long term people will start making archives of AI mistakes and train the AI on those to try to make them less common.
Try it on clip-front: https://rom1504.github.io/clip-retrieval/?back=https%3A%2F%2...
It handles bad, and it handles anatomy. If there aren't single images that cover that - that's exactly what language embeddings solve for.
...This part admittedly trip me: How is it that a system that mirrors every variation of human creativity dystopic? Human creativity is softly bounded by the environments we interact with & the techniques created/(taught to us) for creating such works, along with the knowledge & philosophies that were also created/(taught to us). Ultimately, human creativity is limited in terms of contextual data. The entire art genre of retrofuturism showcases this intentional lack of data in practice.
Scenario: A HASDMLASKD drive doesn't mean anything at first, until a general guiding focus is given to the concept of this drive. Only when it's been given some context do our imaginations fill in the gaps (e.g. space, or storage, or for submarines).
If there's a system that encompasses/surpasses the area that human creativity exists in, that doesn't mean the "oh woe to humanity" doomerism that comes in quick reaction to such a system. It just means that there's a system that can be systematically learnt from & help augment current creation capabilities that lead to more works in the future: Such doomerism is only warranted within a nihilistic context of "humanity will never surpass X", when a more appropriate "X will help increase the area that humanity lives within" could be slotted in.
I think it's a lack of imagination. It's hard to imagine the jobs of the future. We assume work is a fixed sum game, but given new resources we would take different goals and make different plans. It always depends on what is possible, not what was possible.
Exactly this. I'm personally awful at drawing. I'm also awful at most design software. I'm really bad at bringing a concept that lives in my head into the world. I've noticed that playing with Stable Diffusion has allowed me to create things I wouldn't have been able to otherwise. It allows me to create art for projects that I otherwise would've made a lame logo for, or used some stock photography. I don't have the money to hire a skilled artist anyway, so this gives me new possibilities.
First they should sort out the legal question of training the AI on copyrighted material and propose use cases that the general public will find value in. Censorship can be dealt with later.
The request was blocked because "violence was detected". It was a hand-drawn image of a video game boss attacking others with a scythe (it looked like this: https://www.mobygames.com/game/myth-ii-soulblighter/cover-ar...)... there wasn't even any gore, he was mid-swing. I'm a 50 year old guy, I'm not a 10 year old boy, and I don't need to have "violence" (seriously? a hand drawn painting of a video game scene??) censored from me. This nanny-state upstream censorship is BS... I'm just a nostalgic old nerd who wanted a UWQHD version of this image, for sentimental reasons, and this misguided rule stopped my joy.
There is absolutely no evidence that plain nudity, nor hand-drawn violence (which has pervaded comics, video games and movies for decades) has a detrimental effect on human psychology. And yet... the Puritan influence still exists!
At least, if I ran SD 1.5 locally, I could render whatever I wanted to again, but now I can no longer do even that if I use the 2.0 model. This is dumb. Apparently, I'm a "freak" for thinking this.
Things are changing so fast it feels better to just wait until we’re no longer in this phase of having to relearn the tool constantly. In other tech getting in early is important to keep up - with AI generators, I feel the promise is that as the tech gets better, you’ll need to know less and less to use it .
User experience is a big part of image generation. Yet, Midjourney 4 can output better images with easier prompts.
The user perception of OS is mostly third party software availability and driver support. I've used all three, and as far as "operating system" comes, Linux is the only thing that feels good to use.
But, I suppose the analogy isn't entirely inaccurate either. MJ is entirely owned by someone else, it can be removed at any time, including anyone's availability. While SD is flexible, allows a stable foundation to be built on, and you can extend it any way you want...
I couldn't have built a gRPC based wrapper around MJ, set up a bot that listens to prompts sent through telegram, and post back images. Hm.. or I suppose one could do the same with with MJ API... so, bad example :D
People in SD subreddit have been finetuning the SD-models, so depending what you want to do, it should be doable.