I am hoping that AI art tends towards a modular approach, where generating a character, setting, style, and camera movement each happens in its own step. It doesn’t make sense to describe everything at once and hope you like what you get.
At the very least image generators should output layers, I think the style component is already possible with the img2img models.