Where we see shapes, AI sees textures
quantamagazine.org
quantamagazine.org
"AI" can detect shapes just fine; simply use a different dataset and you'll get a different result with the same algorithm. I mean, just look at the classic MNIST: it's all shape, no texture, and neural nets work great.
"...it is observed that while the minimum pattern that could be distinguished from the unwrinkled reference surfaces had a wavelength of 760 nm, the amplitude of this pattern was only 13 nm.
This shows unambiguously that the human finger, with its coarse fingerprint structure in the sub-millimeter range, is capable of dynamically detecting surface structures many orders of magnitude smaller and indicates that nanotechnology may well have a role to play in haptics and tactile perception."
[1] Feeling Small: Exploring the Tactile Perception Limits https://www.nature.com/articles/srep02617
But the insight is that although both features are informative, classical training approaches result in networks that over-fit on texture information and ignore shape information.
Table 2 of the discussed paper finds that joint training with "Stylized ImageNet" (IN images after a style transfer) and classic IN followed by a fine-tuning pass on IN results in improved classification accuracy over IN training alone.
The former are easier to learn for multiple reasons: The nature of higher frequency data means that you have more samples.
(There are far more furry textures in a picture of a cat than there are toes, paws, legs etc.)
Low abstraction data forms are also learnt in earlier layers of the network - more abstract forms are resolved through the understanding of how less abstract forms group together, making the low abstraction forms a prerequisite to more in depth learning.
When it comes to "shape", part of the problem is definitional. In the MNIST example, the network will fairly easily learn features that are nodes, like where two lines intersect, but it has trouble with paths, like "an unbroken line that smoothly intersects itself." CNNs have trouble distinguishing things like a spiral from a set of concentric circles because locally they look the same.
So does the human visual system; see e.g. https://en.wikipedia.org/wiki/Fraser_spiral_illusion .
The spiral vs. circles I agree with.
TL;DR; Logical reasoning isn't based on a "prepondernace of informative." Human-like image-recognition seems, to me, to quire both sort-of statistical reasoning, which depends on the preponderance of information and logical reasoning, which doesn't (not directly).
That's seems sort-of true but "Informative" has a number of different implication. It seems pretty inevitable that a huge image dataset is going to carry the most data in the form of textures.
The thing, a human being can (often) extract considerable meaning from an image of a stick figure or a black and white outline of a big cat, neither of which involves that much data. Effectively, human image recognition is able to both use texture-qualities as a rough guestimate for the source of an image but also look at the overall image structure and see "differences that make a difference". I mean, I can look at a puzzle piece and determine it probably goes in particular area but that's only one facility, "knowing" what an image "is", involves more.
I think MNIST is too simple to have much bearing on the problem. MNIST shapes are two-dimensional; actual object classification requires 3D shapes (and all possible projections thereof), which is a much harder problem.
Since this is just a message board discussion, my wild conjecture would be that the model has no idea about the existence of the third dimension, or basic physical concepts like lighting and distance. And without that larger world model in which to situate things, texture is maybe the easiest thing to cling to.
Actually I don't think humans vision is based only on shapes or textures. There are many novel shapes of planes and we know that's a plane no matter what. We use understanding, with our knowledge of the world, not only the information coming from the eyes. I think it is much larger than what is encoded in a deep learning neural network.
I'm not well-informed about the current state of visual recognition DL, perhaps someone who is can tell us more about whether that approach makes sense.
For example https://www.researchgate.net/figure/Visualization-of-example..., where you can see (somewhat, if you zoom in) that layer 1 neurons are interested in very simple features, like strong horizontal edges, or particular gradients.
Lighting changes indicating edges
Gradients indicating a smooth curve
Texture maps that discern between matte, smooth, translucent, often in depth enough to determine the texture material itself.
Reflectivity
Changes in reflectivity when reorienting oneself or the object
Binocular Vision
Parallax effect between frames
A large database matching objects to memories and experiences
Many other techniques for tracking and seeing movement.
In the above, Parallax and Binocular vision, and consulting one's memory are the only ways to determine the sizes of the objects being looked at.
It's why those optical illusions tiny/big rooms work so easily; neither we nor CV algorithms would be able to discern the sizes of objects without prior memory of such objects or by walking around it.
Understanding shapes/outlines is useful for manipulating objects and avoiding obstacles.
Maybe the object recognition develops a bit later and builds on top of how brains already understand the world.
These networks are correlation-finding machines, and will simply latch onto the simplest correlation possible that produces the correct results. Only by explicitly controlling the correlations in the training data can you force the network to focus on the attributes that you wish.
This is why domain-randomization (such as the approach that NVIDIA is taking) helps support generalization -- it's a very direct and efficient form of regularization, removing correlations that do not generalize, forcing the network to work harder to find correlations that do generalize.
This also has the added benefit of being something that we can tie back to requirements and physical properties, an important ingredient in making these systems understood and safe.
TLDR: The use of texture is inherent to the ImageNet dataset and not to deep learning / ConvNet. Training on less-textured versions of ImageNet drives the ConvNet to focus more on shape.
If we want more human like machine vision then having passes in image processing that deal with more abstract data sounds like a great idea
https://github.com/tensorflow/magenta/blob/master/magenta/mo...
Doesn't quite seem like a break through insight to me?