Somebody posted a nice comparison between DALL-E 2, Mid Journey, and Stable Diffusion: https://twitter.com/fabianstelzer/status/1561019187451011074.
Nevermind because it showed this weakness in the model in understanding what a dollar bill is. The most novel result being the image where the baby's visage appears inside the dollar.
Regarding Abbey Road, I found it interesting that the model's concept of a public person spans their lifetime evidenced by the images where the contemporary images are used. Also interesting to me is the model's weakness in understanding specific people.
Then again though, I haven't been clicking on every DALL-E post so maybe this is old news.
I had imagined that some parts of the training data would consist of the actual image, and it would find a good match for it somewhere deep into it's artificial consciousness.
If it created something very close to the original it would be overfitting. So it flails around rather randomly in photos-of-Beatles+Zebra+Abbey Road space. The results are novel, but they're also monstrous, distorted, and unartistic.
Says who? Maybe you don’t like the style, maybe it is derivative, but it is definitely artistic, as you allude to by saying it is novel. Monstrous and distorted can be a style.
“mathematical art, 1924, litography, abstract generative art” generates https://pbs.twimg.com/media/FanUkREXEAEH6su.jpg — Can you pick the influences? I can’t.
"low poly game asset, Cthulhu monster, 2000 video game, isometric view" generates https://pbs.twimg.com/media/FanTf3IXgAETece.jpg
I thought https://pbs.twimg.com/media/FanNq6sXkAENyU5.jpg was really interesting, because the LEGO logo is not a 100% faithful reproduction even though the LEGO logo is so ubiquitous - showing that the images are generated.
From a longer tweet thread by @fabianstelzer comparing DALL-E 2 vs Midjourney vs StableDiffusion: https://threadreaderapp.com/thread/1561019187451011074.html
I don’t know why it’s got the over-18 warning. It’s 100% SFW.
Perhaps as a blogger I am extra salty when relatively low effort stuff gets upvoted over things that take a lot of work to write (in some cases, those being things I have written). But hey, that's life.
The only cover that really worked from my point of view was the Velvet Underground one, and perhaps the Rolling Stones one. Abbey Road came closer than what I thought it would, but was pretty bad ultimately, and the other three really had nothing usable.
Thus, I look forward to the AI-generated versions of famous works that deepfake the original cast into speaking (hopefully) better-written dialogue. Imagine when this technology is widespread, fanfic authors rendering their interpretation of works with the descendants of DALL-E. Everyone gets the dream adaptations and sequels and finales they want.
And if everyone stacks the cast thusly, there are no opportunities for new actors to work with old and be mentored + opportunity to simply works since every bit of work is what adds up to make a journeyperson.