Adversarial Examples That Fool Both Computer Vision and Time-Limited Humans
arxiv.org
arxiv.org
I opened the paper hoping to see some examples of images that look to me like one thing on first glance, and something else on closer inspection. The best image is the one with the spider on a blurred-snake background, and that's not going to trick anyone who looks at it for more than a second.
The humans were shown each image for either 63ms or 71ms. That's 1-2 frames of a movie. So whilst the result is important, it's not as surprising as you might expect.
The point is: the process that humans go through when they learn is completely different than the process that contemporary neural nets go through. No one has yet come up with a theory that combines all of the features of human learning into an implementable algorithm. It will surely happen eventually, but there are at least a few more conceptual breakthroughs that will need to happen. Minor tweaks to back-propagation won't do it.
But I think "feature detectors" are exactly what the earlier comment was referring to, e.g. a Gabor-wavelet-style decomposition of the retinal image. Deep learning systems have to learn those; we're born with them.
Well, that's one theory. But I think it will turn out to be a lot more complex than that. One thing that I haven't seen anyone pay much attention to is feature detectors in the time domain, which we clearly have. We notice movement as a fundamental feature. Our movement detectors can actually be triggered by static images [1]. One of the ways we distinguish dogs from cats (I believe) is by the way they move. It would be a very interesting experiment to use CGI to make a dog move like a cat and vice versa and see how those are perceived.
It makes more sense to interpret the comment as saying that humans don't learn an internal image representation. Humans do learn representations of bridges, aircraft, cats, etc. But those are built on top of an image processing/representation system that we are born with, analogous to raster graphics?
Edit: Maybe I'm misreading the comment. What's definitely built in at birth is things like edge and orientation detectors. A zebra detector would be a surprise.
Actually, that is what I meant, though I don't have a reference for cat detectors per se. But there is ample evidence for innate feature detectors of comparable complexity (e.g. human facial expressions), even if the actual target is something other than cats.
Indeed, the look of visual hallucination "form constants" is intimately related to the machinery in the visual cortex that detects edges/contours/surfaces: https://www.math.uh.edu/~dynamics/reprints/papers/nc.pdf and such visual hallucinations can be elicited without any drugs or physiological interventions/defects -- just via diffuse flickering light: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3182860/
i was trying to build a bit of intuition about this paper by considering a trivial case where there are two classes, and the classifier is linear, in very low dimension. consider the following trivial example:
https://en.wikipedia.org/wiki/Linear_classifier#/media/File:...
in this example, the decision boundary from different classifiers is shown as H_1, H_2, H_3. The "universal" perturbation for each of the three classifiers would be a small vector normal to the each classifier's decision boundary. This paper defines "universal" perturbation with respect to the choice of input from the population of inputs, but each "universal" perturbation is optimised specifically to target a single model (aka classifier).
both H_1 and H_2 do a reasonable job of separating the two classes, but the H_1 decision boundary with the smaller margin is more vulnerable to misclassification if inputs are perturbed by a small vector normal to its decision boundary.
You can imagine translating the input space a bit to the right -- this would result in H_1 misclassifying say 2 out of the 18 data points shown, whereas H_2 (the SVM generated decision boundary with maximal margin) classifier would not experience any errors.
I may have this wrong but it seems the authors are very interested in a particular subset of classification errors common to machines and humans.
What is that subsets of "transferable" mistakes are trying to find? What does the existence of a particular subset of "adversarial example" (sneakily doctored images) tell us either about human or machine brains?
1. The global reach of networked telecommunications permits a small quantity of sociopathic murderers to operate from beyond jurisdictions that can reach them, and also assists in destroying evidence of their interference.
2. Computational power enables force multiplication, such that even just one sociopathic murderer could exploit software flaws across millions of vehicles, simultaneously.
3. Some software exploits will work against self driving cars, which could never work against an ordinary person, and of course, vice versa, but not so much via remote control at a distance, when people are the operators, while we still lack electronic interfaces to our central nervous system.
We were discussing "tricking" the computer vision component of a self-driving car, not getting software access. That's still a concern, but it's an entirely different set of security requirements that we already face in planes and existing cars.
>painting fake lines on the road, or dropping a cutout of a pedestrian onto a highway
sounds like
>"tricking" the computer vision component of a self-driving car, not getting software access.
Humans can obviously be tricked, in a variety of ways. But adversarial images take advantage of the fact that image-recognizing neural networks do not fit their image recognition into a full fledged understanding of the world like we do. So a few pixels here and there can make a truck look like a panda and the algorithm never says, "But wait, pandas are mostly black and white and this is mostly yellow," or, "But I don't see legs anywhere, or ears."
Optical illusions mostly don't cause high level image misclassifications. To the extent that they are anything similar, they're the reverse: using our general world understanding to cause glitches in our information processing, such as cases where you think something is darker or lighter than it is, or bigger or smaller, or bent or straight. Those are your mind applying rules that are based on "how the world usually appears at a high level" to an image where those rules do not in fact apply.
Cached responses, for example. A question in a text form, which includes a clause pointing at a flaw in usual answer, gets the same usual answer.
Effectiveness of pointing and calling technique [0], which forces full conscious attention.
And other results in psychology, which can be explained by artificial neural network level dumbness of some behaviors. (If 7 workers can build 7 cars in 7 days, then how many days would it take 5 workers to build 5 cars? Is Winnie-the-Pooh a boar or a pig?)