They also just released v5.1, which seems to be quite a bit better than v5.
https://paperswithcode.com/sota/text-to-image-generation-on-...
Parti and Imagen are still on top, followed by Dall-E 2.
If their model is so great, why are they afraid of benchmarks?
For a while I was using an FID variant for evaluation during training, but didn't find it very helpful vs just looking at output images.