The model was trained exclusively on stable diffusion generated images, so can be unpredictable with non-generated images, specially images with people in it
Agreed, results not good for this style of images. I have a model training on a much bigger dataset of image-prompt pairs which should perform better on this.
The model underneath this is trained only on data from SD 1.4/5 so this would be expected. I have another model training which covers all models you mentioned which should perform well on those
True, images generated through some UIs have prompts in meta data, aim here is to work on images people find online with no metadata. So it doesn't try to read the metadata but actually predict a similar prompt
It is based on an image-captioning model, so a different approach then CLIP interrogator, though you are correct that aim is not to get the exact prompt back but actually get a prompt to generate similar styles of images
It works by using an image-captioning model finetuned on SD prompts, so it may be outputting common known artists based on their occurrence in the training data
It actually works on top of an image captioning model, SD takes in keywords as well like "artstation" and "octane render" which are not covered in standard captioning so that is why the difference between using an off-the-shelf captioning model vs this
Hey all! Quick colab to try out Img2Prompt, get accurate prompts from stable diffusion generated images. Works with stable diffusion v1.4/5 images for now