I haven't read the paper, but my guess is that the 50K images in the real-world epoch are not just real images from the cifar-10 dataset, they're 50K random images from cifar-5m. I'm also guessing they don't ever compare performance between a model trained on cifar-10 vs. cifar-5m, they only compare performance of real vs. ideal. So in effect, you can ignore the cifar-10 dataset.