Ten Years of Image Synthesis
zentralwerkstatt.org
zentralwerkstatt.org
You list reasons why text-to-image is easier than unconditional image gen - but isn't that the point? They're both under the umbrella of image generation, which is ultimately pushed forward.
"Interestingly, unlike models trained on ImageNet, where training tends to collapse without heavy regularization, the models trained on JFT-300M remain stable over many hundreds of thousands of iterations. This suggests that moving beyond ImageNet to larger datasets may partially alleviate GAN stability issues."
This suggests GAN will work even better on, say, LAION-5B. It's just that nobody tried.
"GANs are not flexible enough to be effectively used in text-to-image generation" is an absurd statement. GANs have latents, there is no reason why just feeding CLIP to GANs shouldn't work. Again, nobody tried because most GAN works preceded CLIP.
We should try GANs again. It was abandoned without any evidence whatsoever that diffusion model etc is required. GANs have genuine advantages over diffusion models, like much faster sampling.
[1]: https://arxiv.org/abs/2105.05233 [2]: https://www.gwern.net/GANs
The key to AI art is that while the universe is infinitely complex, the visual form of most things humans can recognize follow a collection of natural patterns. Anything humans can do can be computed, the only issue is being able to set up the scope for the patterns we seek to replicate. For games like chess, a bit easier, for images that represent label-able things, harder but still manageable. The main problem of making things beyond this is giving a valid scope to the whole problem: what patterns, precisely, do we intend to replicate?
After a few dozen prompts, you quickly learn how the image generation is simultaneously amazing and dumb. The dumb part is the next 10 years of improvements.
This is like Google 1999.
https://arxiv-sanity-lite.com/?q=diffusion+image+editing&ran...