GigaGAN: Large-Scale GAN for Text-to-Image Synthesis
mingukkang.github.io
mingukkang.github.io
Nice to see GANs making a resurgence, as they are much faster and more controllable. Hands seem to still be messed up though. Are hands just a dataset issue or is there something really hard about it? My first thought is that faces would seem to be more difficult and that has already been mostly solved.
I find it a bit hard to believe the model is completely incapable of rendering these prompts, unless the training data was strictly filtered to exclude artist names and their styles as much as possible.
1) Can I run this, say, on Colab? I see no code.
2) Any interactive open source tool to visually explore the latent space of a model’s representations for any given prompt, or for faces, buildings, etc.? Like they exist for fonts [1].