Here's an example: Stressful Shapes
Dall-E: https://i.imgur.com/JBkSh0y.png
Midjourney: https://i.imgur.com/C02Zq3i.png
On the other hand, here's a specific prompt: "nerdy yellow duck reading a magical book full of spells"
Dall-E: https://i.imgur.com/FMKZ8zc.png
Midjourney: https://i.imgur.com/lpsg6af.png
See https://imgur.com/gallery/U5zJMcU
Comparison of two prompts, "poorly futuristic landscape by a 5 year-old" and "poorly drawnn highly detailed futuristic landscape dotted by mahcinery and tall buildings by a 5 year-old"
Also, https://imgur.com/gallery/jvEClos
Comparison of "poorly drawn red sports car in the street of a city by a 5 year-old"
Edit: forgot about crayons :D
What puzzles me is if the Getty Images logo can sometimes appear. If you only have a Getty account, you get rid of the logo and can legally use them royalty free?
And I don't see the king of Belgium anywhere, two pictures have absolutely nothing to do with the prompt (no king, no speech, no audience, no cucumber), one has the speech and audience but no king or cucumber. Graphically, they are deep into the uncanny valley. Only the third image is kind of right, if you really stretch your imagination.
- - -
Not sure if this is why, but with OpenAI’s Dall-E, you can’t use public figures. You can use proxies, such as “60 year old banker with salt and pepper hair” and then fill in the rest, e.g. “handsome 60 year old banker with salt and pepper hair giving a speech while standing above 12 cucumbers”:
https://i.imgur.com/qYKOWM1.jpg
Telling it oil painting can fudge who the person is, then pick one that’s close and generate variations:
https://i.imgur.com/QRbV7aM.jpg
Or use a reasonable photo and then use edit and in-painting to try to improve the implausible subject. This takes a photo from the first prompt above, erases the lower half of image, and makes a new prompt for the lower half, while keeping just enough of the upper half to orient the collage, e.g. “[photo_edit] + banker giving a speech to cucumbers bin full of cucumbers”:
https://i.imgur.com/OmOK1HF.jpg
- - -
Over on MidJourney, where it’s happy to use public figures so long as you’re not violating terms of service about their use, first a couple prompt experiments with King Philippe of the Belgians.
https://i.imgur.com/KkgIz2w.jpg
https://i.imgur.com/ekf9ypG.jpg
Then one upsized plausible painting from among those, where the actual command was “King Philippe of Belgium talking in a large group of cucumbers --q 2 --uplight” which is pretty basic.
Very insightful tip on how to harness the "creativity" of Dall-E and the like.
I see how the phrase "king of belgium" was too vague for Dall-E, so it didn't produce anything recognizable - but changing the words into known details, like "banker" and "salt and pepper hair", worked effectively to generate concrete imagery.
Hilarious results. :)
> Dall-E: https://i.imgur.com/FMKZ8zc.png
How well it learned all the common prejudices!
"nerdy" == wears glasses
I'm applauding.
I'm looking already forward to AGI based on the current approaches… It will lead us finally into a better world, for sure. /s
Isn't that great? The world will become a better place with AI everywhere.
We need especially more AI in law enforcement, and such…
AI should make important decisions. Because it bears the same prejudices as humans. So it can replace humans just great. ;-)
Now, if 'criminal' rendered as a black male 90% of the time rather than a crouched white male wearing a cheesy burglar mask and a sack over his shoulder, then I could see your point about perpetuating prejudice rather than stereotypes.
Let's not confuse the AI with "buts", just say that he is giving the speech to cucumbers.
Lastly, specify some style, because this would probably not work out as a photo.
My single try is not bad at all and it could definitely be improved.
On the other hand, it is more difficult to get it to produce absurd results like these.
my prompt: King Philippe I. of Belgium giving a speech surrounded by [[[[large green vertical cucumbers]]]], digital art in the style of Greg Rutkowski
It's a bit risky to invest too much time because every generator is different and they change the underlying model frequently (see the beta of MidJourney yesterday), but if you do it for passion or curiosity there is no problem.
Now I'm experimenting with a local installation of Stable Diffusion (well, not really "local" because I have an old computer) and the prompt is only one of the things you can tweak. There are num_inference_steps, guidance_scale and other parameters.
For simple prompts with little additional guidance, all the diffusion image generators I've seen/used will produce output about like what the author linked most of the time. There are always a few gems, and honing in via prompt engineering helps immensely.
With Dalle-2 I get a satisfying result in >50% of attempts and I'm a beginner.
With Midjourney the result almost always looks great, but often misses some part of what I wanted. I'd say Stable Diffusion is similar. The results are seldom crap, but it's difficult to bend it to produce unusual situations. And in SD it's difficult to keep the entire objects in the frame, but that's a different problem.
On top of that DALL-E2 has generally issues with anything dealing with multiple objects. A single person will render fine, groups of people will generally give artifacts. Attributes will also be spread across all objects in the scenes, not just the ones you specified in your prompt, so doing anything more complex will require manual uncropping und inpainting, not just a single prompt.
Anyway, if you avoid the obvious weak spots and holes in the training set, DALL-E2 output is for most part pretty amazing out of the box. It's really more a top 50% than a top 1%.
The biggest bias when it comes to published DALL-E2 images are the prompts. Most prompts you see online are not the actual prompts, but funny descriptions made by a human after the fact. The actual prompt are often much longer and sometimes completely different.
Perhaps this rewrite may yield better results:
"King of Belgium gives a speech to an audience of cucumbers"
And from my experience getting high-quality output from AIs takes a bit of finesse. Not quite unlike crafting a good Google query
so... yes
A twitter user figured out which words they were using by generating a lot of images with the starting prompt "A sign being held that says "
Yes, people tend to share the best of the best. However these results seem especially bad, like bottom 10% bad.
Also: my sense is that getting the best results often requires a lot of extra coaching with style/detail words. As we can't see the prompt here, we don't know what sort of style/details were requested. GIGO.
https://labs.openai.com/s/9YF5WxF1GoZVdzpLBAQYp2Zg
https://labs.openai.com/s/wBfHevs9hIZXvzkJ686mFmn3
https://labs.openai.com/s/M0i029fZnYjQXHFobUpw7eun (best of batch imo)
https://labs.openai.com/s/Hf4z0M9M3KBr6IaaKsjEt9Mx
A few more tries didn't manage to create any photorealistic shots with actual cucumbers-in-seats – perhaps due to the absurd contrasts required – but shifting to a 'cartoon' style with the prompt "editorial cartoon of the King of Belgium giving a speech to many cheering cucumbers, professional illustrator" got a lot closer:
https://labs.openai.com/s/4IonSKYkl0okhNvzJmEAH30K (good)
https://labs.openai.com/s/ZeadCzZ9WqeASYXPlOb13wDV (good)
https://labs.openai.com/s/nERf6bALKEBsQQBvPVsAH7o4 (good)
https://labs.openai.com/s/7dsZu3bwtfGZZxxJYtg9lTVf
If I had more time & credits to burn, I suspect working off those could eventually hit something really apt... but it takes some work & tinkering.
There's one screenshot showing the prompt – but in $CURRENT_YEAR, I view all screenshots with at least a little suspicion, especially when there was a way to highlight the pseuod-watermarked image – OpenAI's native 'Share' – that would've provided stronger proof, direct from OpenAI, of exactly the prompt associated with an image. Hoaxes are everywhere! I've added DALL-E bottom-right color-squares to non-DALL-E images, & seen others do the same, as a subtle joke.
So I generally believe the OP, but don't rule-out the possibility there's been tampering to make some point.