That said, you can sort of get a default look and feel if you just give a short prompt, and then it will tend toward the ones that are favored by RLHF. I prefer very long prompts.... as long as it will allow.
So if you do "a cool treehouse" you'll get sort of the default look. It will be very different if you say "treehouse, naturally occurring, in an old beautiful tree with branches that are low and spread widely and have lots of character and hanging moss and thick bark and curvy roots and mushrooms on a rocky outcropping from a mountainside. photograph, golden hour, sun through trees, damp from rain. Treehouse is part of tree, with fractal forms and live shaped wood and stone and stained glass and glowiness. art nouveau, gorgeous colors and fantasy design"
It's funny that the same people who complain that AI is "cheating" and uncreative, often are the ones who go to so little effort to get good results. It's not like it takes any arcane knowledge to get good images, but if you can use some imagination and string a lot of descriptive words together you can get so much better results.
Furthermore, i found it's very easy to tweak the general design. A cool image, copy the prompt, tell it to make it more X with Y and Z, and you start producing a really neat prompt.
So far as someone with a lack of mind-image but who enjoys creating computer graphics (3d, animation mostly) it's proven as a really neat test bed. Hallucinations are almost a feature in this to me, granted these aren't strictly that - just saying i find it's RNG flavor over my prompt is really nice for exploring.
Sidenote, i entered your text - looks great!
Yours does too. Very different feel than mine, but beautiful.
So my hunch is that humans prefer high contrast and high saturation images!
Most feedback processes for generative models are based on asking the user to draw immediate sentiments rather than having them provide deeper art and style critiques
50mm (optimally with "Nikon" or "Canon" or similar) or 35mm will probably get you the most "natural" looking FoV. (Adding "lower" # fstops will get you a lot more depth-of-field/bokeh.)
The old adage is "f/8 and be there" so f/8 might get the most natural images if you want to specify an fstop(?)
This is where things like img2img in StableDiffusion really came in handy as you could simply apply an entire prompt like a photoshop filter..
I've got two LG 27" 4k monitors, same model number but produced several years apart, and while one monitor can easily show light grays like #EEE, it just looks white in the other.
The paper: https://arxiv.org/abs/2207.12598, of course using CFG change the sample distribution from the training distribution giving it that specific look.
Basically on A/B tests, humans tend to prefer more saturated, "punchy" images. Which is also why iPhones tend to do the same thing.
For artificial images people also seem to prefer stylized "dramatic" styles as well.
And the model was finetuned to match.