That is interesting. My immediate hypothesis would still be that these effects and artifacts are due to biases in the training data. A group of WH40K figures is probably more likely to be identified as a "squad" of Space Marines than just "space marines." Likewise, I suspect there are probably many more pictures of Stonehenge up on the internet than there are of McDonalds'. Or, maybe the training process just somehow found Stonehenge more interesting than McDonalds (I couldn't blame it lol).
One thing I've noticed is that sometimes, the exact same prompt will generate very different images. Once, I got an image of a blue sports car, off to the side of the road, a kind of craftsman-ish looking interior of a house, and two pictures of some random outdoor forest type area. I don't remember the prompt I used, unfortunately, because I ended up writing it off as a glitch, but it would be interesting to see if anybody else has had it happen.
> words that translate your idea into the latent space well,
So, this phrase gave me a sort of a thought: what if certain words or phrases, like, say, "McDonald's" and "Stonehenge," are just so far apart in the model space (again, likely due to biases in the training set, or just the fact that there aren't many McDonald's' restaurants at Stonehenge), that the more interesting or common or unusual one serves as a kind of attractor and dominates the generation process most of the time?
Do you know if these effects are documented anywhere in the prompt engineering literature?