It's quite important to emphasise the dichotomy between the 2 current approaches in the image synthesis today. 1. Implicit distribution learning - VAE or autoregressive based techniques 2. Explicit learning of the distribution - GAN based models. The way these two model the distribution is fundamentally different.
Fundamentally, GAN presents huge drawbacks when it comes to the actual inference during the synthesis. There have been dozens of models with workarounds but most of these present new challenges on their own especially instability and mode collapse being one of the primary.
VQVAE2 as the most advanced VAE based technique has eliminated major drawbacks of VAE and GAN and has produced phenomenal quality [1]
However the main challenge in the area is not synthesising just any kind of image. VQVAE2 is doing that already very well. Where none of the current techniques win today, is the multi-object image synthesis. That requires a new paradigm in the architecture and distribution learning.