A New Lens on Understanding Generalization in Deep Learning
ai.googleblog.com
ai.googleblog.com
I don't think so, Tim. They observe identical performance between cifar-10 and cifar-5m, because the generative model for cifar-5m learned to replicate the distribution that cifar-10 was sampled from. It's the same dataset.
What's surprising is that Google has an enormous dataset called JFT where they could have tested this without this confound. Just shrink the images and you can make something cifar-like and something cifar-5m-like
Also, you use K (thousands) and $K$ (latex K) interchangeably; it's really hard to decipher is K is a variable or what you mean.
I haven't read the paper, but my guess is that the 50K images in the real-world epoch are not just real images from the cifar-10 dataset, they're 50K random images from cifar-5m. I'm also guessing they don't ever compare performance between a model trained on cifar-10 vs. cifar-5m, they only compare performance of real vs. ideal. So in effect, you can ignore the cifar-10 dataset.
It seems like they should have downsampled from an actually large dataset rather than generatively upsampled from a small dataset. Unless I'm missing something?
That being said, I think it would have been much better if they compared some non-convolutional architectures just as a sanity check
Edit: after I wrote this, I checked, and ViT-b/4 is actually a transformer architecture, not CNN. So they did this! And it stayed very close to the same error range from ideal as the CNNs. I am much more confident now that what they did is fine
To clarify some other comments on this post: In all settings, we compare "Real World" and "Ideal World" for the same underlying distribution. Eg, we never compare CIFAR-10 and CIFAR-5m, we only compare "Real World CIFAR-5m" vs "Ideal World CIFAR-5m".
The CIFAR-5m result is interesting, but it's misleading to the reader since the dataset is so contrived. And yet you lead with it on your front page. There's wayyyy too much hype going on here.
Will be interesting to replicate those results on different data.
Overall, one idea is experiments on data sets that have structures but where neural network don't generalize well. Or data sets where neural generalize even better than "the real world".
But it is defined: In the real world, you have a finite dataset, and so have to use it over and over during training. You can shuffle the order, augment it with transformations, divide it into batches, but ultimately you reuse the whole dataset (other than the subset withheld for evaluation) in each training epoch. If the trained network deals correctly with never-before-seen examples (whether they were examples witheld from the training set or brand new ones), then it has generalized well.
Ideally, you wouldn't have to go through any of that rigmarole with a finite dataset, every sample you train on would be fresh and unique, and you would have an endless stream of them. That's the 'ideal world' scenario. The same standard for generalization applies, except that every training sample is never-before-seen as well (so, no epochs).
I have no idea how "you have a finite dataset" defines "real world generalization". This line of reasoning seems completely incoherent.
Anyway, while "real world vs. ideal world training" is defined, you are correct that generalization is not, at least here. But definitions do exist, and more or less conform to what I wrote. Withholding a subset of training data for validation is meant as a proxy for the entirely new samples that will (presumably) be encountered upon deploying the model, but if the training data is biased or otherwise not representative of the real world data, then it is probable that testing the withheld data won't reveal the problem. In other words, the model generalizes well within the parameters of the training set, but not to the real-world data that has a different distribution, because the training set wasn't representative.
There is some interesting discussion of this within Gwern's investigation of the apocryphal 'tank story':