That is to say, if there is some text to image generation program, and the human generates 1 image based on the first phrase they thought of, that image they generated is probably not going to be very interesting, at least based on the image generation stuff I've played around with so far. It often takes time and several attempts / variations on the phrasing to get something which is interesting and novel.
People have been finding tricks like appending the words "Unreal Engine" onto the phrase in order to produce certain types of results. Some people actually go into the code to tweak things to get desired results.
I do wonder how true this will be going forward. It feels like this media synthesis stuff is progressing at the speed of light, and perhaps it'll only take 1 iteration / attempt to craft the interesting and novel stuff that we're looking for.
So at least at the moment it's maybe not so different from anything else. I think of it as exploring the possibilities in a given space. I work on video game development, and it's the same story there. If you're creating a game prototype, you might have an initial idea which sparks development, but once the initial seed is implemented the game often takes on and is guided by a life of it's own, shaped by the possibilities in that given space which you are narrowing down into a subset which seem to work and gel together. That's how I view it anyway.