Measuring the Tendency of CNNs to Learn Surface Statistical Regularities
arxiv.org
arxiv.org
However the bigger insight is that deep networks aren't learning high level features. This paper has very strong evidence that deep networks are just more fancy statistical models as opposed to developing more abstract representations. When if you make network more deeper, you are just developing more higher capacity models as opposed to more higher level features. This is quite damning... If you were hoping current form of deep learning will open the door to AGI, your hopes should be shattered by now.
Some of the things to investigate are (1) can proper regularization help turn this around (for example, forcing gradients to be as small as possible) (2) are there any other architectures using gradient descent that might actually develop abstract features.
This is definitely crisis time in the field and therefore more interesting than ever :).
For instance, there was a talk at NIPS this year where the researcher developed virtual stickers that can be added to a photo and cause the image to be misclassified (eg you have a "banana" sticker that you can apply to any picture to make the classifier classify that picture as a banana). The underlying issue, IMO, is that the classifier has only a limited number of categories from which to choose, and can't develop new categories (such as "adversarially perturbed picture"). The modified image does look more like a banana than anything else, and as the classifier never trained on adversarial images as a separate category it doesn't really make sense to make any other classifications.
For humans, we just happen to somewhat understand the image that is being used against us. How is this different?
There are various semantic features of an image and we presume that our visual system is extracting and operating on these semantic features. So it seems important to ask whether the features causing the misclassification can be represented in a semantic domain of the image. Shapes, (relative) sizes, gradients, high level patterns, etc are examples of semantic domains. There seems to be a distinction between the kinds of optical illusions humans are susceptible to and the kinds CNNs are. In the graffiti illusions, the placement of the image that causes our 3D object recognition system to kick in can be described through angles, lines, focal points, focal lengths, etc. These are all semantic features of the scene depicted in an image. Contrast this with CNN adversarial perturbations which have zero semantic features that we recognize. This seems like an important result. It means that our visual system is robust against certain classes of adversarial images, namely the small imperceptible deltas that trip up CNNs. To trip up our visual system requires certain combinations of semantic input which are harder to exploit (larger delta is harder to exploit).
I get your point to an extent. However, I think it is presented stronger than it is. Specifically, the semantics of imagery are not much more than non frequency analysis of the pixel data your eye sees.
It's hard to be certain because we don't know the details of our visual system.
This would be true even if there were intermediate representations being learned.
http://www.eggie5.com/129-Paper-Review-Measuring-the-tendenc...
... we feel the need to stress that it is not fair to compare the generalization performance of CNN to a human being. In contrast to a CNN, a human being is exposed to an incredibly diverse range of lighting conditions, viewpoint variations, occlusions, among a myriad of other factors.
Why is it that evolutionary heritage is so rarely stated as an important factor in these comparisons (as far as I have seen)? I appreciate that evolution can be framed as another form of learning, but it is much more powerful than that employed by standard CNNs, in that it can change the network structure (and the I/O interface) as well as the network weights.
I don't see how this follows. What we're not focused on is blurred, yes, but what we're focused on is at least partially related to our ability to recognize patterns (e.g. reading from peripheral vision is very difficult).
The fact that CNNs are tricked by imperceptible perturbations while the semantic content is held constant is highly informative information, don't reject it for superficial reasons.
The standard fair of CNN+RELU+Residuals are very powerful modelling tools but that also means they're prone to model degenerate regularities if they exist. This paper shows that they do exist and that these models are picking up on them at least to some significant extent.
Object detection from peripheral vision is not difficult though. I think you're overestimating how much of our vision is actually clear.
>The fact that CNNs are tricked by imperceptible perturbations while the semantic content is held constant is highly informative information, don't reject it for superficial reasons.
Yes, but that is not this paper. These perturbations are not imperceptible, and unlike other adversarial examples the model adapted well when allowed to train on them.
Also, it looks like the model did reasonable well on the random filtered versions, only failing on the blurred versions. The random filtered images looked much more corrupted to me than the blurred ones, which is consistent with blurred images being part of the training for the human visual system, but not randomly filtered ones.
Because we understand high level semantic information, which is largely robust against blurring.
>The random filtered images looked much more corrupted to me than the blurred ones, which is consistent with blurred images being part of the training for the human visual system,
It's also consistent with the neural network training on surface level regularities, which a random filter would corrupt at a much lower rate than a blur filter would.
But only when trained on it. The models trained on only one modified domain didn't generalize well to the others. Although the model trained on all datasets together generalized quite well to the corresponding test sets, that doesn't mean that there isn't some other kind of modification that can fool it.
This is what the authors say:
Our last set of experiments involves training on the fully augmented training set, which now enjoys a variance of its Fourier image statistics. We note that this sort of data augmentation was able to close the generalization gap. However, we stress that it is doubtful that this sort of data augmentation scheme is sufficient to enable a machine learning model to truly learn the semantic concepts present in a dataset. Rather this sort of data augmentation scheme is analogous to adversarial training [40, 9]: there is a nontrivial regularization benefit, but it is not a solution to the underlying problem of not learning high level semantic concepts, nor do we aim to present it as such.
To say that we compensate for the blurring is quite different than saying we can do better (than CNNs?) with blurred images because most of the visual field is blurred. Did you mean to say that we are good at visual recognition despite most of the visual field being blurred? If so, I do not think it throws much light on what is happening in CNNs.