Why Does This Horrifying Woman Keep Appearing in AI-Generated Images?
vice.com
vice.com
https://twitter.com/supercomposite/status/156716492990339481... https://twitter.com/sheslostheplot/status/156737091948789350...
It's also interesting to consider what this image was originally "far" away from.
There are a couple of good followup threads exploring how such a thing might come about in the "latent space" of the model:
The nonsense words of someone who really wants to be the creator of the next big creepypasta.
Or they're lying and this is just a stunt to generate buzz.
It took them about 24 hours to come up with a merch store[1] for this character, too.
[1] https://twitter.com/supercomposite/status/156756520182584525...
That is definitely a tall statement short on evidence. We expect there to be countless islands in latent space - this is basically a related concept to grandmother neurons, although instead of "Jennifer Aniston neuron" we have this "Loab neuron" which has no particular obvious origin. But I don't think there is anything privileged or special about this cluster.
The URL should probably link to the twitter thread in question, the articles adds pretty much nothing and removes tens of images and commentary by the creator.
I even signed up recently just so I could click links and they immediately locked my account for being a bot.
But you are spot on in noting that Vice is engaging in clickbait and misinformation. The tweets are purposefully being creepy because that is the art style of cryptids. But as your second link clearly shows, people aren't going to randomly find images like this (and let's be real, I wouldn't recognize this as the same person unless you told me they were). Stop engaging in melodrama Vice. You can't do that if you're claiming to report news.
what's the point in telling a fish to stop swimming
The article broadly describes a series of prompts, but do we have enough information figure out which AI engine was used, reverse engineer some likely prompts, and try to produce similar results (not exactly the same, as that may not be possible with AI prompts)?
Is it even possible to ask an AI image generator to "produce the opposite" of a prompt?
Is this just an RTFM moment (for me)? or is "producing the opposite" a misunderstanding of how weights work. I only have experience in midjourney, but my understanding is that with midjouney you can weight various prompts as ratios. For example I can build up a prompt to generate an image that is 2 parts "autumn landscape" and 3 parts "birthday cake". But with ratios, isn’t it true that negative weights just discarding that prompt? They don’t produce “the opposite”, right?
To me, this looks like midjourney prompts. At least we can say, that is valid midjourney syntax for sure, and it is using weights. but probably not to the effect the author portrays.
for example, if I prompt "/imagine autumn landscape::2 birthday cake::3", I get this image[1].
But now if I tweak the prompt to "/imagine autumn landscape::-0.5 birthday cake::3", I get this image [2].
Critically, there is no trace of "opposite of an autumn landscape" in this image. It's all birthday cake... 100%.
Indeed, this lines up with some midjouney documentation[3]. A negative weight will try to remove the thing in the prompt.
BUT: what happens when we only have one prompt, and we use a weight to negate it in midjourney. i.e: "/imagine Brando::-1". At least for me, it get this error:
"Invalid parameter
The sum of all of the prompt weights must be positive"
So I'm inclined to conclude that Supercomposite's post is more an act of creative story telling, than an accurate portrayal of their interactions with Midjourney.
But I am still left wondering if there is a true "opposite of" operator in any AI image generator.
[1]: https://cdn.discordapp.com/attachments/1006400576067739749/1...
[2]: https://cdn.discordapp.com/attachments/1006400576067739749/1...
[3]: https://midjourney.gitbook.io/docs/imagine-parameters#prompt...
Thinking in terms of classifiers, an image is almost never categorized as exactly 0% something, instead it's a positive value. Negative weights would make the net optimize for the smallest possible percentage. For the birthday cake, the negative weight is not strong enough to favor any "anti-autumn-landscape" patterns, only to remove features associated with autumn landscapes. But for weights that are all negative, it's plausible that the system will produce all the features that, in the training data, are anticorrelated with the prompt.
https://twitter.com/mattskala/status/1567300206969982979?t=C...
So as someone who works in generative modeling I'll give my best guess as to what is done and what is happening. It is a guess because they don't say everything, but there are some hints. Scambier linked these two twitter threads[0][1] which can give us some insight.
> I'll explain negative prompt weights, in case you don't know. With these, instead of creating an image of the text prompt, the AI tries to make the image look as different from the prompt as possible.
What's important here is that the machine doesn't actually know what the opposite is. In fact, I would ague we don't either. What you can do is use Lp distances from a latent representation. This is where things start to make sense. If in that we find faces as a large distance away, it is also unsurprising that we find many different facial characteristics. These first images look like there is a high mixture between strong masculine features and strong feminine features. These are not things we typically see in reality and combine with our hyperactive brains for recognizing other human faces, we enter the uncanny valley.
Next I don't know if this was done on purpose or not, but there are very clear issues with scene lighting. I can totally believe that this is not on purpose because this is something generators are bad at already. So we have shadows cast along the face in unnatural ways. Upping the creepiness factors.
Now we need to look at important features for recognizing faces: eyes, mouth, and nose. You may have noticed that text to image generators are typically really bad at these. Generators are also typically bad at facial symmetry (why we're trying to get transformers in, but this still isn't working to the degree we would like). In fact, I actually find it more interesting that these are coherent given the explanation of how the latent representation was created.
So I think we have good explanations as to why this would turn creepy very fast. Especially given the hype and that the creator is leaning into it. But these are my best guesses. I can't really know without seeing what is done.
But honestly, I am super interested and would like to see these latent representations and play around with them. This could be a good thing to investigate if you are trying to determine how smooth the latent manifold is, which is extremely important if we're going to make deeper content contributions and rely less on our prompt engineering. Maybe I'll have to play with some negative prompts (if I can find the time lol).
[0] https://twitter.com/supercomposite/status/156716228808747008...
[1]https://twitter.com/sheslostheplot/status/156737091948789350...
My second reaction was to look for articles that don't average out young faces and the result gets even closer. [1]
The traditional simple averaging process removes most wrinkles, rosacea and blemishes because those differ between individuals but the facial proportions match Loab well.
It appears the AI when negatively weighted gives the most average possible result that still matches the concept of a face while picking the most unpopular levels of skin texture and lighting.
It's basically Courtney Cox at 70 after a ten-year whiskey bender.
[0] https://petapixel.com/2011/02/11/average-faces-of-women-in-4...
[1] https://www.researchgate.net/figure/Average-faces-of-young-2...
> “I can't confirm or deny which model it is for various reasons unfortunately! But I can confirm Loab exists in multiple image-generation AI models,” Supercomposite told Motherboard.
this is almost certainly a creepypasta which Vice is for some reporting as if it is real.
He probably found a creepy picture and then started using it as a seed or something like that.
I'll admit I was expecting the "why" (aside from SCP...) to be something like "turns out DALL-E gets live humans and corpses mixed up and in macabre images the live humans get a bit more corpse-y due to this".
https://www.teddit.net/r/nextfuckinglevel/comments/x6d3c3/ai...
I guess it's useful to have an accurate video simulation so you can decide whether it's for you.
The fluidity of the generation seems way more "natural" than most CGI.
Definitely the way forward for every horror video game.
It’s probably because it’s not fully AI-generated. It’s essentially a real video with a filter applied to it[1]
1: https://teddit.net/r/nextfuckinglevel/comments/x6d3c3/ai_gen...
1: https://teddit.net/r/nextfuckinglevel/comments/x6d3c3/ai_gen...
DALL-E itself was Clippy-like, blue, and bean shaped with eyes and mouth more expressive than a Zuckerberg VR avatar.
As for the family? There was none. DALL-E rendered itself on a pure black background
I left out the bit how DALL-E also thinks maybe it's a honey badger, or a lemur with all black eyes, or some sort of sharp wood demon.
The blue one is the one that haunts me because it seems to be the most plausible
This reminds me of playing around with Craiyon a few weeks back. It seemed to think of itself as a coyote who loves soccer. Of course, the training datasets aren't going to have data on dalle or craiyon, so those prompts are going to be more based on the other words, with some randomness
https://www.salvador-dali.org/en/museums/dali-theatre-museum...
Dall-E and other Image generation models are not conscious, and they aren't even intelligent. Stop anthropomorphizing them, it's not helpful. We will likely face this problem for real in the coming decades but there's no sense doing so with current models.
IF I understand it correctly :)
Asking things to draw themselves is fun, I could just as easily ask an elephant holding a paint brush to paint itself and then enjoy the outcome. That's pretty much all there is to it.
Well it's not totally made up. There is definitely something there, there is a real difference between the experience of being a rock and being a human.
The idea then of a P-zombie or some other version of a major intelligence operating with the lights off internally really is spooky.
>Asking things to draw themselves is fun, I could just as easily ask an elephant holding a paint brush to paint itself and then enjoy the outcome. That's pretty much all there is to it.
Agreed. But asking Dalle or whatever model to draw "The meaning of life" and thinking there is some kind of enlightenment in what it draws is ridiculous.
"Consciousness" can be a (fairly vague) term encompassing concretely realizable things like train of thought, the ability to introspect those thoughts, an internal model of self, etc. Humans have these whereas a rock does not, and there's nothing in particular preventing AI from eventually having these.
"Consciousness" as something incorporeal that a physically identical being could lack is nonsense territory IMO (but already heavily debated by people far smarter than me). I don't see how whatever we mean by consciousness can make no physical difference when we're directly aware of it and talking about it in the physical world.
That seems like a surprisingly deep philosophical statement on society in general. And IMHO, along the same lines, "reliably prevent harmful results" is not something that should be pursued to exhaustion.
If you want to talk about statements on society in general, I think the fact that nudity is seen as more harmful than gore, is a bigger problem than an AI generating either one.
In fact, the dataset used to train it received precisely this criticism on HN back when it was originally announced.
Somone here on HN has said that they use these AI models to explore the latent space of the human imagination. I would say that sounds exactly correct.
Bizarre. It's like saying that you are using the newly invented Bicycle to explore the concept of General Relativity. Woefully insufficient.
EDIT: I'm not kidding.
Perhaps there is some thing the machines can see, but our psychology protects us from?
Just scroll through https://lexica.art/ which catalogs 10M+ StableDiffusion images, most aren't even close to being disturbing and are in fact aesthetically quite pleasing.
It's like going to the grocery, seeing a bin of beautiful apples and saying "every apple on the tree is beautiful". You can't draw that conclusion because those apples aren't representative of all apples because they went through a filtering process. In this case some one choosing to sell them. Upthread, some one choosing to share them.
Fun fact, apples grown for market get some level of sun shading to keep them pretty. Too much sun damages their skin and can give them that sort of leathery look. Which now that I think about it is a filter itself, hardly germane anymore though.
Interesting so I plugged in for prompts from “Captain Kirk”, and none of the images look much like William Shatner. More like a cross of an old Chris Pine with a tiny bit of Shatner thrown in..
https://lexica.art/?q=Captain+Kirk
Weird stuff
https://twitter.com/supercomposite/status/156766407375967846...
> To clarify for the press (many are asking for an explanation without jargon): I have brought a real IRL demon to life. Research has found that demons are real and live inside of computers. Computers are like little houses for demons and church is like a big house for angels.
What should we do about the energy crisis? Make them go out and chop wood.
What should we do about quiet quitting? Send them to the gulags.
What do we do about dissidents and people who disagree with us? Send them to the furthest gulags.
What about the poor wheat and corn production this year? Kill all the birds.