Can someone expand on this? I've never heard of this before, at least not in the general case.
Can someone expand on this? I've never heard of this before, at least not in the general case.
They shuffled the labels on their datasets, so there can’t possibly be anything to learn, yet got zero training loss, meaning the network must be severely overfitting. Yet the same network trained with the actual labels shows quite good generalization. So the usual intuition about overfitting and the bias-variance tradeoff doesn’t seem to apply.
When you have nothing to learn, you need to memorize the data. But when there is structure, it is easier to memorize the structure, so the network will learn this first (and will memorize after).
“Understanding Deep Learning Requires REMEMBERING Generalization”
https://calculatedcontent.com/2018/04/01/rethinking-or-remem...
https://arxiv.org/abs/1710.09553
We can understand this using the traditional theory of Statistical Mechanics of Generalization
Briefly, shuffling the labels corresponds to decreasing the effective load on the Neural Network, which pushes the system into the spin glass phase