I’ve been comparing Dall-E, MidJourney, and StableDiffusion. Goes to show how much training set and implementation choices matter. But in all cases, you have to think of the underlying labeled text-to-image sets as paint colors to mix, and prepare a palette accordingly. Still haven’t figured out how to get what I want, but to your point, one can get closer.
- - -
Not sure if this is why, but with OpenAI’s Dall-E, you can’t use public figures. You can use proxies, such as “60 year old banker with salt and pepper hair” and then fill in the rest, e.g. “handsome 60 year old banker with salt and pepper hair giving a speech while standing above 12 cucumbers”:
https://i.imgur.com/qYKOWM1.jpg
Telling it oil painting can fudge who the person is, then pick one that’s close and generate variations:
https://i.imgur.com/QRbV7aM.jpg
Or use a reasonable photo and then use edit and in-painting to try to improve the implausible subject. This takes a photo from the first prompt above, erases the lower half of image, and makes a new prompt for the lower half, while keeping just enough of the upper half to orient the collage, e.g. “[photo_edit] + banker giving a speech to cucumbers bin full of cucumbers”:
https://i.imgur.com/OmOK1HF.jpg
- - -
Over on MidJourney, where it’s happy to use public figures so long as you’re not violating terms of service about their use, first a couple prompt experiments with King Philippe of the Belgians.
https://i.imgur.com/KkgIz2w.jpg
https://i.imgur.com/ekf9ypG.jpg
Then one upsized plausible painting from among those, where the actual command was “King Philippe of Belgium talking in a large group of cucumbers --q 2 --uplight” which is pretty basic.
https://i.imgur.com/QWUaNFv.jpg