> without some additional prompting or fine tuning that encourages it to do something else.
That tuning has been done for all major current models, I think? Certainly, early image generation models _did_ have issues in this direction.
EDIT: If you think about it, it's clear that this is necessary; a model which only ever produces the average/most likely thing based on its training dataset will produce extremely boring and misleading output (and the problem will compound as its output gets fed into other models...).