It seems like the major cost here is that you need to train a high quality StyleGAN for your dataset -- which is very data hungry - at least thousands of images for high quality results. (The famous face generator used millions of faces)
I think that the idea is that you use transfer learning to bootstrap from a general gan to a gan for your domain - and then use that to generate odd data.
Last I checked, transfer learning for StyleGAN2 still produced somewhat wonky results if the datasets were not extremely close. Have there been improvements in this field?
I don't know - but I guess that must be in the paper. I had intended to read this properly after I looked at it when the story came out, but I haven't had time, sorry.
I think I'd seen StyleGAN2 using like hundreds of examples! And I'm pretty sure Flickr-Faces-High-Quality only uses 70k examples.
But that’s not bad compared to the cost of pixel-wise labeling to get training data for segmentation.