Img2Prompt – Get prompts from stable diffusion generated images
img2prompt.io
img2prompt.io
> a monkey plushie on a white background, photograph taken by steve buscemi from a zoom lens, studio lighting, ultrarealistic
Can someone tell me what Steve Buscemi is doing here?
FirstName LastName being the name of a politician.
Thank you.
It’s good to describe a picture but it’s not reverse engineering. The predicted prompt usually has very little in common with the actual prompt. And it’s worse when you use embeddings or fine tuned models.
But as you say, even so, it still has little in common with the original prompt.
This tool is useful for reverse-engineering prompts from the kind of images I want, then generating new ones in the same style.
Very cool.
Also I laughed out loud after putting a selfie into it and getting "mark zuckerberg's face reflected in a mirror, close up, realistic photo, medium shot, dslr, 4k, detailed"
Of course, this is not actually how latent space works. It's the AI's understanding of concepts, not the inherent nature of concepts; that's why every model has its own version of "latent space". Though the understanding of latent space in the story is internally consistent; given a superintelligent image generator, you could do prompt engineering like this.
I often use terms like "sexy", "risque" etc. in the process of getting images that are quite sensible (like military people playing chess). I use img2img repeatedly looking for particular photo-film aesthetics and tend to accumulate prompts. Anyway, this would open me to charges of sexism (or worse "misogyny"), and makes me uneasy about using SD.
Edit/OH: it generates prompts for like Excel screenshots but not for images made with the img2img model at hugginface. Fascinating.
It made me so agitated I made a gallery[0]. Granted, some of those images are strange, but others are just normal people doing normal things.
[0]: https://publish.obsidian.md/zero-chroma-infinity/Image+galle...
Tried it on last image I generated on: https://blazordiffusion.com/artifacts/50/50418_studio-ghibli...
Original Prompt:
> Studio ghibli, rocket explosion, jungle, solar, green technology, optimist future
> 8k, Bokeh effect, Cinematic Lighting, Octane Render, Iridescence, Vibrant
> by Beeple, Asher Brown Durand, Dan Mumford, Greg Rutkowski, WLOP
Img2Prompt:
> a vehicle in the grass, colorful light dust, cinematic lighting, trending on artstation, ultra detailed, art by akihito yoshida
Looks like a decent image classifier, but not useful for extracting the original stable diffusion prompt.
It's not prompt based intended to generate another one, but rather an accessibility tool.
And some related videos:
Seeing AI 2016 Prototype - A Microsoft research project - https://youtu.be/R2mC-NUAmMk
Seeing AI: Making the visual world more accessible - https://youtu.be/DybczED-GKE
https://i.imgur.com/GpGG0SL.jpg
And got this prompt:
> the man with the stupid face of a homeless person, portrait photography, 1 9 7 0 s, street photo, old photography, highly detailed, hyperrealistic
The actual description is that she was Mary Ann Bevan (1874 - 1934) also known as Rose Wilmot, a woman who claimed the title of the ugliest woman in London as she suffered from acromegaly.
Thank you.
I made a few generic military guys for a side project and the actual prompt isn't too far from what I used.
> john cena walking on stage at television talk show, very coherent!!!!!!!!!!!!!!!!!!!!!!
And it described my profile picture on Mastodon as "my husband from the future that looks similar to travis scott and mark owen. he is also a good boy, very!!!, and beautiful!!!"...
Not sure how to take that ;)