You list reasons why text-to-image is easier than unconditional image gen - but isn't that the point? They're both under the umbrella of image generation, which is ultimately pushed forward.
"Interestingly, unlike models trained on ImageNet, where training tends to collapse without heavy regularization, the models trained on JFT-300M remain stable over many hundreds of thousands of iterations. This suggests that moving beyond ImageNet to larger datasets may partially alleviate GAN stability issues."
This suggests GAN will work even better on, say, LAION-5B. It's just that nobody tried.
"GANs are not flexible enough to be effectively used in text-to-image generation" is an absurd statement. GANs have latents, there is no reason why just feeding CLIP to GANs shouldn't work. Again, nobody tried because most GAN works preceded CLIP.
We should try GANs again. It was abandoned without any evidence whatsoever that diffusion model etc is required. GANs have genuine advantages over diffusion models, like much faster sampling.
[1]: https://arxiv.org/abs/2105.05233 [2]: https://www.gwern.net/GANs