Okay so the diffusion model generates image assets, and you iterate on those, and then another model turns your favorites into code, is that right?
It's a lot like having an architect create plans for you before handing it off to a builder. In my (obv biased) experience, you end up getting better/more creative results with this approach. You're using the best model for the job at each specific task, ie a diffusion model as the designer, and a LLM as the engineer.