I want to see a few iterations of describing an image with AI, generating it, describing it again, generating it... Like when passing a piece of text through Google translate back and forth.
It needs a better text to image model, I think. Maybe you can fork it and improve?