> Usually, the image the model creates doesn’t exist in its training data - it’s new - but because of the training process, the most influential images are the most visually similar ones, especially in the details.
Would be cool if this were true, but I don't think it is, because the prompt you used and the captions on the training images are being completely ignored. If two different words tend to be used in captions for very visually similar images, and you use just one of those words in your inference prompt, I'm pretty sure the images that were captioned with the word you used are much more "influential" on your output than the images that were captioned with the word you didn't use. (Like, "equestrian" vs "mountie" or "cowboy" or something.)