Understanding deep learning requires re-thinking generalization
blog.acolyer.org
blog.acolyer.org
"As the authors succinctly put it, “Deep neural networks easily fit random labels.” Here are three key observations from this first experiment:
-The effective capacity of neural networks is sufficient for memorising the entire data set.
-Even optimisation on random labels remains easy. In fact, training time increases by only a small constant factor compared with training on the true labels.
-Randomising labels is solely a data transformation, leaving all other properties of the learning problem unchanged."
And conclusion
" This situation poses a conceptual challenge to statistical learning theory as traditional measures of model complexity struggle to explain the generalization ability of large artificial neural networks. We argue that we have yet to discover a precise formal measure under which these enormous models are simple. Another insight resulting from our experiments is that optimization continues to be empirically easy even if the resulting model does not generalize. This shows that the reasons for why optimization is empirically easy must be different from the true cause of generalization. "
This paper was pretty hyped when it came out for seeming to discuss general properties of deep learning, but the details of it are a little dissapointing - okay so sufficiently big/deep networks can overfit to training data, and that's exciting how?... it's a curious finding, but not one that's all that hard to believe or that is all that informative. Or so it seems to me. I don't see how they justify claiming that they "we show how these traditional approaches fail to explain why large neural networks generalize well in practice."
I suppose the notion is that memorizing random labels implies memorization should also work on non-random labels (and thereby no generalzation to test set is needed), but it seems intuitive that proper labels and gradients with regularization will find the answer that generalizes because that is the steepest optimization path available. I have not read it all that deeply and not in a while, so perhaps their arguments are stronger than it appears to me, though.
a) The traditional view of generalization would argue that neural nets are "too complicated"/have too many parameters/etc to generalize well. And that for generalization you need more limited models.
b) To reconcile this with practical results of CNNs some people tried to argue that while neural nets had a lot of actual parameters, the structure of neural nets reduced the effective capacity of neural nets to not be as big as the parameter counts implied. Same argument for regularization.
c) This paper shows that those arguments are not satisfactory since they really can fit random labels.
It should be not surprising that strategies "memorize all input and interpolate between" are powerful, no matter if as a part of neural networks or other easy-to-overfit techniques (such as random forest).
What's a 'genuine correspondence'? The network has clearly picked out some image features that correspond to the assigned labels. Just because they're not the features you're thinking of doesn't mean they don't exist.
The classifier is, in effect, memorising every element in the training set. It's training a compression algorithm for storing that data.
It should be noted that "training a compression algorithm" isn't always a bad thing per se, because that's how autoencoders work, which is one of the main ways to do deep learning.
The key term in the article is "the effective capacity" of the model. If you have a big enough network, it can simply memorise everything you give it. This makes it difficult to know if such a model will generalise. A much smaller model won't overfit in the same way, but also might not perform as well as a larger, more sophisticated model. The problem in deep learning is nobody can tell how much of the training data has simply been saved somewhere in the model (in an obfuscated and compressed way).
There is some related research about reconstructing the training data from deep networks (which has privacy implications), but I don't have a link handy.
If there is no detectable patters from images to labels, the network does rote memorization. It learns a compact way to remember the label for each image.
This must be the case because the generalisation performance can vary significantly while they all remain unchanged.
Maybe it was just me, but I read an implied "alone" in the sentence you quoted, ie:
"Or in other words: the model, its size, hyperparameters, and the optimiser, alone, cannot explain the generalisation performance of state-of-the-art neural networks."
Indeed a careful hyperparameter choice is the only key now to have good generalization. As I understood it, the goal here is more to show that the correlation between the regularization of the network and its generalization power is far from being clear as it is for other ML algorithms like SVM.
In short, NN hyperparameters help to reach generalization, but cannot "explain" it. It's the key difference here between practice and theory.
Deep networks can generalise to situations where even humans cannot. So the memorizing narrative doesn't survive any scrutiny.
The capacity of a million-sized shallow net might be O(million), but noone's using such a model.
I asked the question on CS stack exchange and no one took issue with the statement that DL had such a large VC dimension. The only counter response was that it didn't matter in practice due to DL's good error scores. But, that still doesn't mean DL is generalizing. Good error is only a necessary condition for generalization, not a sufficient condition.
Perhaps deep networks work well because they learn to memorize "the most common patterns of auto-correlation" they see in the training data at different levels of function composition.
In fact, we do this explicitly in convolutional layers, which by design learn to represent every input sample as a combination of a finite number of fixed-size square filters.
...and the reason why deep networks might be "generalizing" so well is because not all distributions of natural data are equally likely!
In practice, objects with the same or similar labels tend to lie on or close to lower-dimensional manifolds embedded in data space.
...and this concentration of natural data distributions might be a result of the laws of Physics of the universe in which we happen to live: https://arxiv.org/abs/1608.08225
So, yes, deep learning could very well be a really fancy form of memorization.
I hadn't thought of it in this way before. Very interesting :-)
Sure, we can apply even higher level features, and we can generalize from a picture of a cat to a black&white or text description of it.
But we still memorize a lot.
That last bit about "multiple levels" is key. Shallow models like kNN, SVMs, Gaussian Processes, etc. don't do that; they learn to recognize/memorize patterns at only one level.
Second, it's not PR! This stacking of layers, when done right, can overcome the "curse of dimensionality." Shallow models like kNN, SVM, GP, etc. cannot overcome it; they perform poorly as the number of input features increases. For example, k-nearest-neighbors will not work with images that have millions of pixels each.
Third, I'm only scratching the surface here. There's a LOT more to deep learning than just stacking shallow models.
The model, size, hyperparameters, optimiser, etc - all they do is convert the data into a form that can be used to make predictions.
Deep learning is called deep because it is based on multiple levels of features corresponding to a hierarchy of notions. Of course, it can change how generalization as well as other operations are performed but the way generalization is done not a specific feature of deep learning.
The network must not have capacity to hold all the data. It must have a capacity proportional to the number of classes of data (instead of the number of samples).
Another way to arrive at this may be: take a trained network, run it in inference on the training set. Group the nodes of the network into equally sized groups. As inference happens train a smaller corresponding new group of nodes for each previous group by looking only at its inputs and outputs that are exercised. Put the new subnetworks together by looking purely at the edges between the previous subnetwork. The new network is now constructed.
I have not built this. But would something like this work?
That's already known: there are inputs to the network that do not make sense and yet will trigger strong responses. Think of them as inputs that have the same effect on NNs that optical illusions have on the human brain. We infer something that isn't there.
I suspect that as network architectures get better and parameter counts drop these will get harder and harder to construct.