Why do CNNs generalize so poorly to small image transformations?
arxiv.org
arxiv.org
"Do CIFAR-10 Classifiers Generalize to CIFAR-10?" - https://arxiv.org/abs/1806.00451
They use the same procedure used to construct CIFAR-10 to construct a new test set and test a bunch of state of the art results on the new test set.
They see a generalization gap, but the relative order of SoTA results remains (roughly) the same.
So, yes, test set validation is an overestimate of real-world performance of these systems, but progress on test sets is indicative of progress in real-world settings.
And to remember this in context, no-one is asking "do traditional computer vision systems understand images", because they were explicitly just looking at image statistics.
I think this is important point to make and talks to the real life limitations of run of the mill SL techniques. If you can't validate with noisy "real world" data, then you're basically staying inside the box and hoping that you've curated well enough to mimic real world conditions.
I still like RL system for the improvement on this, however it's much more difficult in practice.
> We view this gap as the result of a small distribution shift between the original CIFAR-10 dataset and our new test set. The fact that this gap is large, affects all models, and occurs despite our efforts to replicate the CIFAR-10 creation process is concerning.
> Nevertheless, the accuracy of all models drops by 4 - 15% and the relative increase in error rates is up to 3×. This indicates that current CIFAR-10 classifiers have difficulty generalizing to natural variations in image data.
It remains to be seen if others can think up ways to maintain performance that the authors did not manage.
Those models don't seem to be THAT brittle.
> Our main finding is that CNNs exhibit a tendency to latch onto the Fourier image statistics of the training dataset, sometimes exhibiting up to a 28% generalization gap across the various test sets. Moreover, we observe that significantly increasing the depth of a network has a very marginal impact on closing the aforementioned generalization gap. Thus we provide quantitative evidence supporting the hypothesis that deep CNNs tend to learn surface statistical regularities in the dataset rather than higher-level abstract concepts.
[0] https://www.opticalspy.com/opticals/dr-phil-plait-optical-il...
[1] https://arxiv.org/abs/1710.09829
[2] https://hackernoon.com/what-is-a-capsnet-or-capsule-network-...
I did a quick search and came up with this paper[1], and while they don't use the Fourier-Mellin transform, they do use a log-polar transform[2], where rotation and scaling are transformed to translations. This, they claim, result in much improved classification of rotated and scaled data.
I haven't been keeping up on the latest in the ML field tho, so maybe it's rubbish, but like you I'd be surprised if some extra processing doesn't play a role.
[1]: https://arxiv.org/abs/1709.01889
[2]: https://sthoduka.github.io/imreg_fmt/docs/log-polar-transfor...
How about a head mounted on a ball joint ? Would that count ? (since obviously it would introduce small spatial variations VERY often).
You can make a similar argument for being mounted on a mobile robot (like the human or any other animal body).
Anyway, I think that neural networks are now entering the trough of disillusionment as people begin to discover the limitations. Maybe in the future, somebody will come up with a new machine learning architecture that has better generalization. I'm not expecting gradient descent to give us general AI.
Very interesting consequences for the ML field if my hunch has anything remotely resembling a kernel of truth to it.
Humanly "imperceptible" is a very loaded term. Human perception has billions upon billions of networks worth of filtering going before we even boil down our environment to the "interesting" stuff.
Furthermore, if you take a snapshot of that network after training, you're fit to the training data. The network has lost it's plasticity. Take a potato, put it on the ground, train the network on other potato shots. Now show it a potato shaped asteroid. Now show it a French fry. What is the potato-ness that this potato+detector is ACTUALLY homing in on? Keep in mind, this structure is trained on digital encodings of maps of light and color. The function may not be a perfect semantic detector of potato-ness. It just knows what patterns of bits MIGHT be potatoes. And when you are working on bit level encodings, one bit translates to a lot of change, even if it is imperceptible to a human looking at a rendering on a screen.
Heck, there is no guarantee that the function it's emulating is well defined outside the training data set.
Neural networks are GOING to be fickle. You're trying to coerce "reliable, repeatable. generifiable results" out of a simulation of the same stuff that drives five year olds and emotional people. Consider yourself lucky the program hasn't opened the CD tray and demanded you insert crayons.
The "neural network" machine learning technique is not a simulation of biology. It's a nice marketing phrase. The technique is just math, maybe "inspired" by someone thinking about neurons.
Don't be misled by branding. For a good explanation of why some of these fanciful science terms come about, read Bellman's explanation of why he called his research "dynamic programming".
The crux is that you are programming in such a manner that you can differentiate your function w.r.t. parameters which are then tuned.
The main difference of course being that you are implementing it in silicon rather than carbon, and having no bloody grammar or conception of how in the heck to explain how 'X set of nodes with Y set of weights with Z activation threshold function'=something useful.
They called them Neural networks for a reason. It's been quite a mystery both on the biology (in terms of gray matter) and computing (in terms of ANN) front why it works at all. But it does.
I think I've read Bellman's work before, but I'll take a look. Thanks for the pointer.
Just can't find enough hours in the day to keep on top of the state-of-the-art, and unfortunately the career hasn't budged me in the direction of anywhere I could weasel work on it into my day.
C'est la vie.
This is why the 'neurones' used in ANN's are nothing like biological neurones really - both in the properties of the activation functions and the structure of the net as a whole (which is difficult as we still know relatively few details about the structure of the human brain on a neural circuitry level).
Neural Network architecture (sideways skyscrapers that get more focused as the network plays out) does not incorporate generalization much and this is sort of the point: a specific input data shall generate the resulting target, and any small variance would decidedly result in something different at the other end of the NN. Unless there were some way to incorporate a broad pattern on the data as it comes in... For example, run the NN, the NN shows results and also produces a fingerprint hash, now when you put slightly mal-aligned data into the NN, also provide the previous fingerprint result, and now (maybe) we can derive a device that will result in the same product and similar products within a range of the input data, to satisfy the case that the fingerprint stays true even when the data is slightly different.
Just some ideas on the topic, if we could solve pattern generalization where the viewing window did not have to line up perfectly with the perceived pattern, that would be very powerful; we would have the ability to measure unaligned states, but an algorithm that approaches this power also must exploit some natural phenomenon, such as quantum superposition, if we are to complete our iterations in reasonable time. Without some sort of preprocessing step, such as sorting (that can ensure algorithms run quickly) it seems difficult to try and make a real neural net that is significantly more than a very specific compressed archive.
What's really interesting is that as a stochastic process we could both generate two separate, functional, result-bearing neural networks (weights and layers) that achieved same or similar results while being completely different at their basis (in number of layers and the actual activation weights). And right now, there's no way to determine [elegantly] how similar our neural nets are. So, perhaps in some ways, the ability to compare neural network guts meaningfully will lead us to the ability to encode more elegant view ranges on data sets.
Have you ever seen kids learn the tables of multiplication ? Or addition (but people don't seem to remember from their past that they learned addition from tables). Doesn't that give you pause when you're saying humans can generalize ? Because clearly, this is not how we teach our children.
Humans don't learn that if 2+2=4, 3+2 must be 5. No. Humans learn that 2+2 is 4, then they learn that 2+3 is 5, then they learn ... and so on and so forth.
Even when you see online players investigate a game. Whether it's chess, zelda, or starcraft. You keep seeing the same thing. Humans don't learn by generalizing. They learn how to act in every situation ... one situation at a time. So it's not just 6 year old kids doing this. The game advances because there's a few specific individuals that spend entire days making nonsensical moves, and on rare occasions they find a new move, see what happens, and tell the world (in chess we're talking something like once a year, in mario it seems possible to do it once every few months. It's scary how much tenacity is required for such a process).
And for the vast majority of humans, it never even gets to that point. They just don't have the tenacity.
So if you think humans are AGI, which you seem to do, then clearly such a state can be achieved with disappointing generalization capabilities.
Not sure what AGI means.
Yeah, you make a strong point that there are axiomatic must-knows for any domain, and generalization would require some level well above -- or some layer well abstracted from -- this axiomatic knowledge.
The thing is the most difficult mistake was 5+3 being 14. I tried teaching them to approach the problem from different angles. 3+5 ? Fine. 4+3 ? Right. 6+3 ? Right. 5+3 ? 14 ... it took hours.
In kids, most mistakes are corrected quickly, but a few persist. That reminded me that I have a few choice mistakes that lasted an incredibly long time despite getting corrected again and again and making no real amount of sense (meaning I know dozens of ways to calculate them, and I know how the numbers progress and I know how to use distribution to solve it, to "move the beans" (5 + 7 = 5 + 5 + 2 = 10 + 2 = 12. They present that logic, doing the tens first, as moving beans around. Which makes sense), and yet, there's a few mistakes that stuck. So you have a lot of time to think while correcting one of these sticky mistakes.
There are a few moves in this direction but without a better framework for computational intelligence, which we can get to faster through reverse engineering, we’re going to keep hitting these walls. This road includes methods for artificial neurogenesis, genetic encoding of the developmental program at the level of neuron and axon driven by artificial gene regulatory networks (AGRN) evolutionary methods to evolve populations, and adversarial ecosystems for open ended evolution. Unfortunately none of these will give you a quick win....
Complexity of biological systems is so mind-blowing that we likely don't even have computational capacity to simulate one of those processes in veracious detail. In addition, we can't even comprehend primitive, 1000-layer deep neural networks, not mentioning significantly more complex neurons (that are performing some local, protein-based computations alongside electric spikes we measure). We should be happy that Deep Learning surprisingly somewhat works, obliterating "classical" ML and make use of it to the full extent of its capabilities. Romantic idea that we can construct AGI with it should be retired.
If you’re a materialist then you’re going to believe that all of the minute biological details are relevant. If you’re a functionalist you can work at a more abstract level. From the 12 years of research I did in this space (both CS and Bio) I can tell you for certain that we can learn a lot more from a functionalist approach.
And the reason we “can’t understand” 1000 layer deep neural networks is again because we don’t have a good theory of these types of computational systems. For instance I can tell you that most of the interactions w_ijk are completely spurious and if you cull the network you actually end up with a circuit that you can understand in the same way any EE undergrad can identify a 3 bit adder from the circuit topology.
1.http://science.sciencemag.org/content/315/5814/961?casa_toke... 2. https://onlinelibrary.wiley.com/doi/full/10.1002/syn.21972 3. https://www.sciencedirect.com/science/article/pii/S000689931... 4. http://onlinelibrary.wiley.com/doi/10.1002/hipo.22365/full 5. https://www.sciencedirect.com/science/article/pii/S092523121... 6. https://www.sciencedirect.com/science/article/pii/S089662731... 7. https://ieeexplore.ieee.org/abstract/document/7124441/ 8. https://www.sciencedirect.com/science/article/pii/S107474271... 9 http://science.sciencemag.org/content/353/6304/1117?casa_tok...
They don't have good local behavior: if a -> f(a) and b -> f(b) are fitted with a higher-degree polynomial then the image of f between a and b will often be infinite (meaning there is some value x between a and b where f(x) is infinite), starting at very high degrees (ie. deep networks) there will probably be many such points.
Intuitively I think of it like this: if you look at the "real world" as a function you can make a couple of very general observations. F(x) -> doesn't have very much information about the world, and it's very hard to make sense of. Most of it just doesn't seem relevant. d/dx F(x) ... much more relevant. d^2/dx F(x) ... also pretty interesting. d^3/dx F(x) less interesting but occasionally important. d^4/dx F(x) ... nobody cares (also if you take camera images and calculate this, it'll be almost exclusively zeroes).
Secondly there are strong "domains" in the real world that we just seem to be unwilling to accept. Polynomials are good in the sense that if you get the equation for a stone dropping onto your foot really, really tighly correctly fitted, that equation holds up for the movement of an entire planet, which is great. But why bother ? It is much more valuable to be able to predict whether a stone will fall on my foot than how Venus will move. That's if you get it right. If you get the polynomial degree of your equation wrong ... it makes utterly ridiculous predictions. Many other approximation methods don't suffer from this problem
Doesn't happen with spline, beziers, even taylor approximations have better behavior.
In machine learning, a convolutional neural network (CNN, or ConvNet) is a class of deep, feed-forward artificial neural networks, most commonly applied to analyzing visual imagery.
(Source: https://en.wikipedia.org/wiki/Convolutional_neural_network)
Perhaps the answer is simple : REM (the awake variant), and https://www.youtube.com/watch?v=quJEyTvDdfY
So how would a network perform w.r.t. translations if it was expanded, i.e., topology similar to the CNN but with all parameters expanded, and trained on translated images?
Pointing first and then classify may help in solving those issues. Maybe?
No shit! An earth shattering result. :P
>While VGG16 has 5 pooling operations in its 16 layers, Resnet50 has only one pooling operation among its 50 intermediate layers and InceptionResnetV2 has only 5 among its intermediate 134 layers.
Aka being computationally cheap with pooling causes weird sampling issues.
The paper has a terrible pompous title that's just wrong. Modern CNNs generalize wonderfully on small translations. They just managed to break a couple of old CNNs and go on to claim they broke AI research or something.