Thanks so much for this! This was an exact problem my team had at a recent hackathon, and we decided to use diffusion instead. We theorised that the initial image parameter might help, but we didn't have time to test it. Documenting this was really helpful. If you have any other articles or advice on using VQGAN (or other generative models) with CLIP (or other language embeddings) I'd love to see them.