But maybe with 5 million real examples sampled from the distribution CIFAR-10 was sampled from they would in fact see a difference. Maybe the generative model is capturing only a limited slice of the diversity that the ideal model would really see.
It seems like they should have downsampled from an actually large dataset rather than generatively upsampled from a small dataset. Unless I'm missing something?