>you're never going to be able to perfectly describe an image in your head. Language is too limited.Pretty much this - and that's why both "AI art panic" and "I won't need artists anymore as I can just prompt anything" (which OP seems to try) are based on the wrong premise.
In general, people tend to make 2 false assumptions:
1. That SD is a ready to use product (it's more of a "middleware" model that is designed to build products upon)
2. That text to image is all it can do, while the most power is in style transfer, finetuning, and the ability to be guided with higher order hints than just text alone.
What would be best right now to use it as a content creation tool, is to combine SD with several other models to have temporal stability and tagged-3D-to-2D, and to use some software toolkit with a pipeline like that:
- quickly layout a mock-up scene in 3D with rough assets, just like in game engine level editors, possibly tagging the geometry or objects with short descriptions, like "middle-aged man", "Volvo semi", "pine tree" etc. No need for detailed geometry, just stick figures and rough shapes.
- using one of the multiple available techniques, train your style on your reference images (which might either be curated output of the model itself, or a specific visual language you constructed)
- enter your prompt, which doesn't need to be complex and convoluted as you already provide the precise hints on what you want
- let the model render the result, based on the depth, tagged geometry, rough rendering providing shadows and reflections etc