If you generate 25x more images, you can afford to cherry-pick.
I'd be curious to see how a vision model would go if it were finetuned to select the best image match to a given criteria.
It's possible that you could do O1 style training to build a final stage auto-cherrypicker.
Better metrics (assuming goal is text->image) would be some sort of inception score or CLIP-based text matching score. These metrics are computable on single samples.