When it's done, it generates a web page with the images and a simple UI for feedback, then uses the feedback I give it to iterate.
It has always been very good at styling things well (i.e. if I'm making infographic-style images, I give it an example of one and it's good at copying the style/fonts/colors/etc.), but I suspect I get better-than-average results in terms of the content just because it has built up so much context from doing this and taking my feedback across so many different brands and images.
The product dimension thing really stopped being an issue for me with ChatGPT Images 2, and the new 2.5 is better. If you're running into trouble, there are some cases where Gemini is better, but they're rare. But I will say if you're prompting yourself and running into issues, you should get Claude/ChatGPT to help you write the prompt that goes to the image model. They're very good at it.
A little more on my process here: https://theautomatedoperator.substack.com/p/the-dumbest-way-...