And this applies to language / code outputs as well.
The number of times I’ve had engineers at my company type out 5 sentences and then expect a complete react webapp.
But what I’ve found in practice is using LLMs to generate the prompt with low-effort human input (eg: thumbs up/down, multiple-choice etc) is quite useful. It generates walls of text, but with metaprompting, that’s kind of the point. With this, I’ve definitely been able to get high ROI out of LLMs. I suspect the same would work for vision output.