We asked 100 humans to draw the DALL·E prompts
surgehq.ai
surgehq.ai
The biggest problem w/ current DALL-E is that it's really really bad at details, unless it's a stereotypical one. DALL-E delivers pretty good impression, but almost always fails to fill in the last mile. This is because DALL-E currently doesn't know what "perfection for human" is like..
Also, often times, DALL-E doesn't have a definitive style, meaning it composes images with different styles in one place. This is quite noticeable in the second image for "Teddy bears, mixing sparkling chemicals as mad scientists". The bubbles are obviously not in the same style. This hardly happens for humans because, welp, when someone draws a picture, one person draws with a limited set of tools in a single style. DALL-E itself is not subject to such limitations.
For example the space ones have incredible style, texture and shading, but the basket ball lines are nonsensical. While the human ones are quite basic and artistically primitive but all of the easy things are flawless.
And then the bears have an eye on their forehead and missing half their mouth mixed with extremely good textures everywhere else.
The human drawn ones were much more interesting to look at, as art, usually because of the emotion and hints of narrative conveyed by facial expressions. I also liked the giant cat.
These days famous artists often don't execute the artwork themselves, they leave that to craftspeople. They have the idea, and the fame to imbue it with attention which connects it with the semantic fabric of the current art scene.
I imagine DALLE-3 will generate at 256px directly, and use something like latent diffusion for a free upscale to 1024px. This should eliminate most of the high frequency artifacts.
With AIs like DALL-E, they usually end up looking terribly uncanny and sometimes all over the place.
In other words, while GPT, DALL-E, etc. are good at making things that superficially match the prompt given to them, vaguely seeming coherent, the moment you inspect them closer they do not guide your attention, or create a broader artefact with inter coherent sub parts, the way a human artist/writer would.
It’s hard to tell if that’s an inherent limitation of approaches based on statistical sampling rather than “true understanding” or not. That criticism has held so far over the last decade of progress, so I doubt “more parameters/training data/resolution/etc” is the answer.
You can buy and shovel out all the Maroon 5 you want but you can't be cool unless you co-opt Rage or GnR or some band that's legitimately cool and meaningful to people and outside the reach of what everyone knows was already for sale. AI art is always gonna be corporate garbage regardless of how perfect it looks because the people who guide it will always be corporate douchebags.
What you're really perceiving is that every detail in the human work has humor. The humor runs through it; everything, mistakes and all, is shot through with what we visually can see as a sense of the personality of the person who drew it. Not that an AI can't imitate each one, but it can only think inside the conceptual box of what it's trained to imitate.
For instance, the first graphic concept that came to my mind about astronauts playing basketball in space with cats would be to have the whole scene of a cat dunking in space shown in the reflection of an astronaut's visor. The raw literal image is cheap crap without the viewer seeing some human perspective that's trying to be conveyed... in other words, engineers need to go to art school because they're misunderstanding the value proposition behind human-made art.
Plus it got the lines on the basketball all wrong
The human ones are basically a collage of items accurately created to a strict specification, while objects in the AI-drawn ones are all integrated and more artistic and creative.
Although this is probably because the humans were people motivated by money who were paid to spend 15-30 minutes to make the artwork and just tried to check off the requirements (I can't possibly imagine a decent artwork being created in 15-30 minutes).
It's very interesting though that humans are more robotic than neural networks!
The modern process of producing music would basically be unrecognizable to anyone 40 years ago — it's completely intertwined with technology, and far more automated. Yet music is as important as ever, and amazing music is being made (will politely side-step the pitfall of debating whether music was better 40 years ago!)
So I'm excited to see how visual artists incorporate tools like Dall-E into their artistic process.
I don't think you'll have many takers here suggesting that things were magically better 40 years ago. (We know about survivorship bias and that there was less of the design space search so much more novelty/amazing).
I do think we've gotten a few more axes of exploration but also have in the mainline of music has gotten much more homogenous in some ways as well, which is kinda sad.
Sophisticated tools are a bit of a trap. People tend to create in ways that their tools make easier. Tastes evolve around what's being created. And tools evolve to match those tastes, which in turn really optimizes everything to a specific local maximum. And tools cause loss of skills and dependency upon the tooling.
Ha, fair point. I must not realize how old I am, because I was attempting to reference the music of the 1960s and 70s, not 1982, which I agree is not many people's idea of the golden year for music ("Come On Eileen" notwithstanding).
> Sophisticated tools are a bit of a trap. People tend to create in ways that their tools make easier.
No doubt. Ableton, logic, and protools have drastically altered the norms of what modern music is "supposed" to sound like (ie tuned vocals, quantized drums etc). I do wonder what the next generation of music tech will bring.
⸻
1. I have a theory that most people tend to favor the pop music of their high school years. No idea if it’s true or not.
I would rank that theory with evolution, relativity, and the germ theory of disease as being as close to true as any theory is ever likely to get.
I think that pretty soon we'll see many hit songs that have a large contribution from AI models. It's possible that the variety can increase as well and new styles can appear, as the "good" generated songs are fed back as training data.
Currently we have youtubers putting out crude AI voice generated videos with basic visuals, and its still funny and entertaining because the story is good. But imagine if anyone with a good idea could start producing studio quality works to accompany their meme video.
For example, the first image has a basketball where the lines don't line up correctly, a visual artifact that lets you immediately know it was an AI. Human artists would have made all the lines line up correctly.
Fourth image it gets a bit confusing because the art is originally abstract, but the astronauts hands are again very easy to spot, the second biggest artifact are the cat heads, they are just close enough to be abstract but there's warping there that no human would do.
Neither of those is the biggest problem with the cat - it has two left front legs and no right front leg.
I thought the biggest tell in the first four images was that the AI makes no attempt to draw correct lines on the basketballs.
Playing basketball? Multiple balls! The planet is a ball!
I actually like this dollop of creativity. Humans tend to ground their pictures in reality more than necessary for prompts that sit firmly in the realm of fantasy.
Well, that picture captioned "DALL·E or Dalí?" is almost certainly drawn by an AI, because most humans have learned that you don't draw weird shapes anywhere near ones crotch, it's just something a half-decent human won't do.
Joke aside, the AI generated pictures are amazing. See those radishes? Most human-drawn radishes are happy. And those generated ones are happy too, even without explicit instruction to tell the AI to do so. I guess it really captured what humans collectively want.
Happy radishes? :)
Anyway, I wonder if things like these cause a feedback cycle that just makes everything super boring. Humans sometimes get bored and do unusual things, what about the machines?
Most human drawn cats are evil. Most AI drawn cats are happy.
It's great and I love it, but the wow factor is fading and the easily recognizable problems remain.
Alternatively, imagine one where every website is flooded with mass-produced bland "art" like Corporate Memphis... we know what SEO spam looks like today, and I've been occasionally seeing what looks like AI-generated images show up in image searches. This feels like the latter --- good enough to look the part at a brief glance, but clearly "off" if you take a closer inspection (see comments about the lines on the basketballs.)
P. S. Has anyone demonstrated actual style transfer with Dall-E? As in, provide an image by one artist and telling it to draw it as another artist would? Something along the lines of "Vincent Van Gogh's self portrait as drawn by Jack Kirby'.
https://www.lesswrong.com/posts/r99tazGiLgzqFX7ka/playing-wi...
So there's definitely pressure on illustrators on upwork, fiverr etc.
That's exactly what DALL·E2 is already capable of. You can paint over a region of the image and tell it what you want in that area:
https://youtu.be/qTgPSKKjfVg?t=77
You can even use this in-painting ability to enlarge images. For example you can feed it this:
https://preview.redd.it/z7y2tbwjtny81.jpg?width=680&auto=web...
And turn it into this:
https://i.imgur.com/Zp4MWXX.jpg
https://i.imgur.com/xsEydvk.jpg
Source: https://old.reddit.com/r/dalle2/comments/umk5vw/science_fict...
And also worth mentioning that each of those iteration only takes like 10sec, not days or weeks as with a real artist.
The one thing DALL·E2 is however missing is being specific, you can tell it to draw SuperMario and it will come up with something that looks like somewhat like SuperMario, but it is taking a lot of creative freedom along the way. As far as I can tell, it isn't able to draw the same character in different situations. Every image is unique and there is so far no way to coax it into drawing a series of multiple images that are consistent with each other, as you would need to illustrate a short story for example. Best example I have seen so far is this, but even that remains very abstract:
https://old.reddit.com/r/dalle2/comments/umrkef/variations_o...
While it goes in the right direction, it's all just a jumbled mess and far less impressive than when Dall-E focuses on a single character, e.g:
https://old.reddit.com/r/dalle2/comments/u5vajy/super_mario_...
https://old.reddit.com/r/dalle2/comments/ubr2xd/a_real_life_...
https://old.reddit.com/r/dalle2/comments/ufzk2f/incredible_d...
It would have been cool to also add the watermark to the non DALL-E images and then do the disclosure of what is what.
For those that don't know, the watermark is the bottom right rainbow squares.