This Image Does Not Exist
thisimagedoesnotexist.com
thisimagedoesnotexist.com
I've been keeping a close eye on DALL-E 2, especially images posted on Twitter and Reddit[0], and the biggest sign that it's a DALL-E 2 image to me are the edges of objects not being as crisp as you'd expect in a human-created image.
The composition that DALL-E 2 creates is typically fantastic, but it isn't perfect when it comes to the edges, details, and textures (fabric, for example).
I still think it's incredible for inspiration — I don't think a typical artist/designer will use DALL-E 2's output as the final image, but rather as a great starting/mid point for their final image. I can see myself using DALL-E 2 to generate something akin to a wireframe that I'll then go in and add the details to manually.
The fact that I'm using the subject matter as much as the image itself to discern AI art is a sign of how good the tech is though. Especially since the subject matter in question is typically "stuff that would be really challenging for an AI to conceptualise" and the image itself (not necessarily the AI's only attempt of course) is a sound attempt at rendering it.
I see most of the utility as the opposite: the artist still needs to come up with the original concept, the overall compositions are unlikely to be better than an artist would conceive, but even in the hands of a non-artist a bit of text can produce "good enough" images (that can be cropped or touched up to remove rough edges) instantaneously which would take a lot of time for a skilled artist to produce.
Counter-example: there was one image I got that was a Gameboy as a concrete brick. That seems like a prompt you’d give DALL-E 2, but it was actually of human creation
It makes sense that someone would think to make that because a gameboy is basically already a grey brick. So it's really just a change of texture- it's a comparison of two things that look similar that a person would probably make but an AI might not. it has some artistic intent.
The other DALL-E prompts are just "a <random subject> doing <random thing> in <random medium>", like "a dog building a fence in watercolor".
Would it make sense for a person to paint a dog building a fence in watercolor? Probably not unless they're writing a children's book.
Of course most human-created images are equally shallow, but if you do see depth then you know it came from a human. Put this test to an art historian/appraiser or architect and they would probably score very well.
For example, if you prompted a human and DALL-E to draw "a fork in an electrical outlet", the human would probably portray the fork as shocking and exploding, whereas DALL-E would probably just portray a fork being put into an electrical outlet and nothing happening. Maybe that's a bad example because there may be enough depictions of forks in electrical outlets for DALL-E to learn to depict them as exploding, but my point is that DALL-E can't reason generally about the composition of the image.
> The composition that DALL-E 2 creates is typically fantastic, but it isn't perfect when it comes to the edges, details, and textures (fabric, for example).
For me the trick for discerning them is asking "what's the purpose of this detail?", namely: if a figure is writing, what are they writing? If a figure is wearing a shirt, does it look like a shirt I've ever seen? Why are they standing the way they are standing? And what material is this statue made of?
For instance, the train at dusk on a bridge was really well composed overall, but some of the bridge cables were slightly transparent.
I think all it would take to "defeat" me here would be greater variety of high-quality image generation AIs.
I expect in a few years that will be the case, unless the "style" of DALL-E-2 is something fundamental, instead of an artifact of its specific implementation.
Also anything mechanical is just way off on AI generated things - a car with the fuel door in the middle of the passenger door, a missing tire but you see no axle and nothing makes mechanical sense.
I can see DALL-E-2 stuff spreading around WhatsApp groups like wildfire, masquerading as something else
Also - please speed up the delay switching to a new image! I bet this would feel way better if the delay was in the 0-0.5s range instead of multiple seconds. (I bet it’s some runaway JS, investigating…)
Edit: `changeImage` is blocked by `await saveDB`, and the POST is 500ing. (e.g. uBlock makes that AWS request fail fast – why desktop people are less affected)
Maybe use `Promise.allSettled` to make them parallel, and prepare the data being sent beforehand to avoid state race conditions :) thanks for leaving sourcemaps on!
I came here to suggest the opposite, as I didn't have time to read the full prompt which generated the image, when it was a generated picture.
So I guess the conclusion is instead: allow users to click a second time on any of the buttons/picture to proceed at their own leisure, so the ones who want to read can read, and the ones who want to speed through it, can do that too.
(Edit: see JacobiX reply, looks like it robbed the concept, possibly because I didn't give it any direction.)
I just prompted GPT-3 (davinci-002 engine on openai) to write a children's book and include illustration suggestions. I didn't even suggest a topic, just said 'children's book'. I fed the illustration suggestions into Midjourney with the qualifier of 'children's book illustration of ...' and pulled the first image of each (it always generates four).
I then asked it for a cover page, last page and a background image for the text.
I assembled the story text written by GPT-3 and images rendered by Midjourney in this google slides preso. The only real creative input i had was the font. I did iterate on the text background image prompt a couple of times to find something that would work:
(Note this link works in incognito if you prefer not to have your google login associated)
https://docs.google.com/presentation/d/e/2PACX-1vT4XWNx2SdEg...
I've been having a lot of fun with this prompt. Part in asterisks is where you set the topic. The last 'Title of the book:' is where you submit, it really helps keeping the generative part on track.
----
Write a four page children's book with the following:
Title of the book:
Description of the cover page with enough detail for an illustrator to draw:
Description of the background imagery used for the pages with text on them:
First page text in quotes:
Description of the opposite page illustration:
Second page text in quotes:
Description of the opposite page illustration:
Third page text in quotes:
Description of the third page illustration:
Fourth page text in quotes:
Description of the fourth page illustration:
Description of the final page:
The subject of the children's book is *how to prevent buffer overflows in c code*.
Title of the book:For this particular case, it seems like GPT-3 is paraphrasing the story ...
Here's what resulted from my prompt above for a children's book about buffer overflows - https://docs.google.com/presentation/d/e/2PACX-1vSp7wc1xCgiv...
Doesn't always work hahaha.
Input:
The subject of the children's book is how to autorotate a helicopter after its engines have failed.
Title of the book:
Output:
"Autorotating a Helicopter"
Description of the cover page with enough detail for an illustrator to draw: The cover will feature a large illustration of a helicopter in the air, with the text "Autorotating a Helicopter" in big, bold letters.
Description of the background imagery used for the pages with text on them: The pages will have a white background with black text.
First page text in quotes: "If you're ever in a helicopter and the engines fail, don't panic! Just autorotate."
Description of the opposite page illustration: The illustration on this page will show a helicopter with its engines on fire, falling towards the ground.
Second page text in quotes: "To autorotate, pull back on the cyclic and maintain a level attitude."
Description of the opposite page illustration: The illustration on this page will show a helicopter with its engines on fire, but the blades are still spinning.
Third page text in quotes: "To land, flare the helicopter just before touchdown."
Description of the third page illustration: The illustration on this page will show a helicopter landing safely on the ground, with the blades still spinning.
Fourth page text in quotes: "Remember, practice makes perfect! So if you ever find yourself in a helicopter with failed engines, you'll be prepared."
Description of the fourth page illustration: The illustration on this page will show a helicopter in the air, with the engines on fire.
Description of the final page: The final page will feature a large illustration of a helicopter landing safely on the ground, with the text "The End" in big, bold letters.
a few possible illustrations - https://ibb.co/c2ZYQps
There is also some kind of artifact on DALL-E's images which makes gradient areas look very low color depth.
Hardest ones to get right are low quality images and images of abstract artwork or the ones that are made in a "smudgy" style already.
It is pretty good overall but I would say that it's almost impossible to fool anyone if the images were in high resolution and you could zoom them in.
It's probably most effective in small images like thumbnails for example.
It's kind of spooky to think how good it will be in a few years and what it will be used for.
There is a long runway of human guided ai generation. Imagine if a single human was able to produce an entire animated TV show and all it took was good story telling skills without requiring 100 animators.
Something I think about a lot is video game creation - what if AI tools let you quickly create 3D objects and a world without a big, dedicated team of artists?
(Check this out by Nvidia! https://m.youtube.com/watch?v=5j8I7V6blqM)
I suspect there was similar wailings and gnashings of teeth when photography was invented. Humans adapt and it's us that decides what to value. We are quite fond of other people so tastes will alter to preserve that balance.
The career path for an artist trying to find something else while AI is developing at this pace is...less clear, and definitely scarier.
Same goes for this really, but perhaps even more so. It's impressive at producing eye-catching photos quality images of simple concepts you might use in advertising campaigns, and can do a good job of reproducing well defined cartoon styles. But the artists working in those sort of roles are already working digitally and using image libraries, and this is a massive productivity boost for them - dozens of "$brand on beach" concept images for no effort at all, and they'll still need all the retouching and layering and colour balancing and combining with other images. And as soon as the brief gets more precise about details of layout and colour scheme or what they want the model to look like, they're either back to their original creative flow or designing text prompts and curating images for the brief is an advanced skill in itself. Non artists can produce "good enough" images too (looks great for memes), but much like non-artists wielding cameras they're mostly satisfying their own whims rather than competing with the business of art
An unpaid intern adding text to AI generated images would be trivial.
It would be nice if they told you in more detail how the algorithm worked. How well curated are the AI and human images? It seems like selecting the absolute best AI ones and the worst human ones would make it easy to make things look better than they really are.
A genuinely "fair" way to do it I think would be this:
1. Randomly collect a few thousand images from DeviantArt
2. Ask people to write a short description of each image, similar to the prompts given to DALL-E
3. Generate DALL-E images from all the prompts
Randomly show a human or AI image with 50% probability.
That should give you a good measure of the actual state of the art.
HHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
and
HHTHTHTTTTHTHTHTHTHTHHHTHTHTHTHTH
are equally as likely outcomes.
To make it more obvious, let's take just 4 flips. There is exactly one combination with 4 heads:
HHHH
Whereas all the following will give you 2 heads and 2 tails:
HHTT HTHT HTTH THHT THTH TTHH
So you're 6x more likely to get 50% heads/tails than 100% heads. Scale that up to 30; the chances of getting 15 heads and 15 tails is not the most common thing but not unusual. The chances of getting all heads is quite rare -- if every human on earth did this test, we'd expect around 7 such instances total.
Overall most AI images are clearly identifiable by lack of fine detail in things like fingers or leaves and a telltail method of blending between colours in low detail areas.
The image definitely does exist. I know they're trying to riff off of "this human does not exist" -- but that site works because the word "human" can mean a person in real life. I do not know of a usage of "image" to mean an artwork that must be painted by a human.
Perhaps giving the "real" art a random crop would level the playing field slightly.
Perhaps have more human generated images, although that might be a limited set and people will just soon remember which images are human made. Impressive though.
The machine generated ones are usually easy to spot if you look closely, because parts of the image are just nonsense (e.g. weird patterns, things out of place, odd focus). Machine generated images clearly have no actual understanding of the real world.
It makes me wonder, are models being unintentionally trained to produce jpg artifacts? How much better would a generator be if it had been trained on, say, pngs?
Some key things that let me to immediate spotting:
- DALL-E 2 is really awful at symmetry, in art pieces it's less obvious because even humans are terrible at replicating features equally.
- A lot of human images are pictures from Unsplash, so they have some pretty recognizable features, in particular framing, shallow depth of field or more commonly, very noticeable ISO noise
There's always a few duds and some definite limits (https://www.lesswrong.com/posts/uKp6tBFStnsvrot5t/what-dall-... )
but on the whole most of the output is very good - the images on this website aren't particularly taxing for it and I don't think cherry picking would be needed to replicate this level of quality.
I think the current post really gives the wrong impression.
Maybe it was cherry picked?
Those 2 I got wrong were heavily stylized.
It's very easy to spot DALL-E artifacts, mostly geometric lines and shaped being distorted in a way that humans wouldn't draw, small details such as fingers and so on being deformed and so on. And it struggles with symmetry.
Looking forward to getting further feedback.
Awesome project!
You might want to consider cropping all the human images to be square as well. It's a dead giveaway when all the AI images are square and the human ones have varying aspect ratios.
All the AI generated images are square which kind of ruined it for me.
I struggled with some of the stylised paintings, but it's clear that DALL-E cannot do detail.