I don't think so, Tim. They observe identical performance between cifar-10 and cifar-5m, because the generative model for cifar-5m learned to replicate the distribution that cifar-10 was sampled from. It's the same dataset.
I don't think so, Tim. They observe identical performance between cifar-10 and cifar-5m, because the generative model for cifar-5m learned to replicate the distribution that cifar-10 was sampled from. It's the same dataset.
To clarify some other comments on this post: In all settings, we compare "Real World" and "Ideal World" for the same underlying distribution. Eg, we never compare CIFAR-10 and CIFAR-5m, we only compare "Real World CIFAR-5m" vs "Ideal World CIFAR-5m".
The CIFAR-5m result is interesting, but it's misleading to the reader since the dataset is so contrived. And yet you lead with it on your front page. There's wayyyy too much hype going on here.
I haven't read the paper, but my guess is that the 50K images in the real-world epoch are not just real images from the cifar-10 dataset, they're 50K random images from cifar-5m. I'm also guessing they don't ever compare performance between a model trained on cifar-10 vs. cifar-5m, they only compare performance of real vs. ideal. So in effect, you can ignore the cifar-10 dataset.
It seems like they should have downsampled from an actually large dataset rather than generatively upsampled from a small dataset. Unless I'm missing something?
That being said, I think it would have been much better if they compared some non-convolutional architectures just as a sanity check
Edit: after I wrote this, I checked, and ViT-b/4 is actually a transformer architecture, not CNN. So they did this! And it stayed very close to the same error range from ideal as the CNNs. I am much more confident now that what they did is fine
What's surprising is that Google has an enormous dataset called JFT where they could have tested this without this confound. Just shrink the images and you can make something cifar-like and something cifar-5m-like
Also, you use K (thousands) and $K$ (latex K) interchangeably; it's really hard to decipher is K is a variable or what you mean.