Hacking the Brain with Adversarial Images
spectrum.ieee.org
spectrum.ieee.org
What would be interesting is if you could establish a correspondence between visual and genetic similarity... neural networks provide a distance metric.
http://www.infinityplus.co.uk/stories/blit.htm
That’s the first one, others are available online too.
rolls eyes I feel like this article totally misses the point about why adversarial ML tactics are interesting. The fundamental reason they work is that computers don't have any abstract notion of a cat, a toaster, a banana, or even gravity or matter. THAT is why they're easily fooled.
To stack up two adversarial images and say, "Look we fooled humans too, OMG humans are just like computers!" is like saying, "We erased the lines on the road and both self-driving cars AND humans got confused, OMG humans == computers!"
I know the article isn't saying humans == computers, but .. come on .. I'm not seeing any merit to this investigation
First of all, when we say that a classifier "misclassifies" an image, we have an agreed-upon definition of what that means: it means we have an image, we have a target label for it and the classifier assigns it a different label.
What exactly does it mean that a human misclassifies an image? More specifically, when the article says that "humans think that they're looking at something they aren't", what in the world does that mean? I mean, I have to assume a confusion between "looking", and "seeing". I can't really believe the article is saying that when I'm looking at the top right image I'm not actually, you know looking at it; but instead ...looking outside the window? Or what?
On the other hand- seeing? Really? Who knows what it is that anyone else is seeing? Who knows what is there to see? Especially when it comes to an image specifically manipulated to be confusing, as opposed to a real-world image? What is the ground truth here, for the image on the top-right? Is it a cat, just because Google reserachers say it's a cat and it's fooling my senses into thinking it's a dog ish? Is it fooling my senses just because the Google researchers say it is? Google researchers are the final arbiters of objective visible reality, now?
To be more precise, who is to say that that image can just be a "cat" or a "dog" and nothineg else? Because, you know, the first thing I thought when I saw that image on the top right was "that looks like a jumbled mess made with bits of an image of a cat and an image of a dog", or something along those lines.
Of course, if you sit me in front of a computer with two buttons, one for "dog" and one for "cat" like in the researchers' setup... well, then you can force me to misclassify the image all you like. But what does that prove? Besides the fact that if you force me to choose one of two things, without knowing what you think is the "right" thing to choose, I'll choose the "wrong" thing a lot of the time?
And it's more than just "making it look like a dog". If you look at it for any length of time it's obvious that it's a picture of a cat with some noise on it. But if you only have 100ms to respond to it you are guaranteed (>95% misidentification) to say that it's a cat. Now what happens if you take this and make your car look like a tree? Someone could crash and that would be terrible! Cancel driving, tear up the roads, ban wheels, cars were a mistake! /s
Human vision is slightly more robust solely because it's had more time to go back and forth with adversaries. Nothing prevents ML from reaching the same levels of safety. Nothing prevents you from deploying attacks against humans that're identical to to the attacks against artificial systems.
I imagine most self-driving car hackers will react to any successes with "holy shit it worked" followed by remorse.
Obviously I can manipulate a picture of a cat to make it look like a dog.
That may be the point, but it's not proven by this study. Cats and dogs have many structural similarities, as do the other adversarial examples (panda / gibbon, cabbage / broccoli). We know that this isn't just an artefact of the human visual system because we know they are very similar in non-visual ways too, i.e. they are genetically and behaviourally similar (compared with random other items in the world, such as bananas and toasters).
This study's choice of images seems to acknowledge that when the human visual system makes mistakes, it does so in a far more robust way than ML-generated models do. Even the fact that this effect is robust across many individuals, versus adversarial ML images being model-specific, demonstrates this. Mistaking a cat for a dog, given a 50ms window, is much less likely to be disadvantageous to us than mistaking a cat for a computer, or a banana for a toaster.
In other words, this study is a long way away from demonstrating that applying a bit of static to an image could make us mistake a car for a tree, whereas in an ML scenario such a mistake seems quite plausible.
>Researchers from Google Brain show that adversarial images can trick both humans and computers, and the implications are scary
So the IEEE is telling me how to feel about this article in addition to presenting the facts. One might even consider that "hacking the brain with adversarial text". :-)
>A worrying possibility is that supernormal stimuli designed to influence human behavior or emotions...
Sooooo, like clickbait headlines? :-)
Brains are very easy to hack. The idea that we're reliable exemplars of rational objectivity and rigorous self-awareness is nonsense.
To start, it should be clear to everyone that it is possible to actually transform a picture of a cat into a picture of a dog. The argument that the author is trying to make is not, in fact, that the human brain is hackable with adversarial images, but that the kinds of strategies which allow you to hack a wide range of deep learning models (to turn cats into dogs) require the introduction of "features" that a human would recognize as being doglike.
The leading image is not evidence of this, however. We have to keep track of the what figures correspond to which claims, because the article jumps around.
The first figure shows us a mask which people agree turns a cat into a dog. This is not evidence of any claim because such masks are guaranteed to exist.
The second figure shows a mask which people don't even detect which causes a computer to claim that a panda is a gibbon. This is basically the null hypothesis of the paper, it demonstrates that you can attack a single computer with a subtle mask.
The third figure shows us an example of an attack that is robust across deep learning models and also to perspective shifts. The argument being defended by this figure is that humans recognize the mistakes that the machine is making as being non-superficial features of laptops and toasters. The author undermines his own argument here, by requiring that the attack be robust to perspective shifts. The point he/she is trying to make is that in order to be robust to many neural nets you have to include non-superficial features, but then convolutes that claim by also requiring the features to be robust to other types of transformation as well. Does the "featuriness" of the attack come from the multiple neural network requirement or for the transformational robustness requirement? The answer is unknown.
Furthermore, I think this argument falls short pretty quickly if you simply hide the computer's answer. Would you guess that the computer guessed laptop for the first image, even if you knew it got it wrong? How far down your list would that have been? Or toaster for the second? The features offered in the second image definitely look like they have some features that are similar to the corner of a rounded metal cube, but this is exactly a superficial image statistic which is robust to change of perspective transformations. My classifier might as well be looking for that exactly. This is not evidence for the authors case.
The fourth and fifth figures are more interesting, but still do not provide any evidence in defense of machine learning.
It's important, for some reason, to point out that the accuracy on image tasks is NOT measured linearly with error rate. 99.9% accuracy is 10 times better than 99% accuracy. The difference between 65 and 75% is about 1000 times less than the difference between 75% and 99% Allow me to submit, without evidence, that humans can differentiate between a cat and a dog with >99% accuracy when given as much time as they need to make a decision about whether or not they are looking at a cat or a dog (with no adversarial mask applied). This means that the difference between human performance on the task without the time constraint, versus human performance on the task with the time constraint, is about 1000 times larger than the difference between human performance on the task with the time constraint and the task with the time constraint and the filter.
This means that if you really want people to fail on the cats versus dogs task, spending time coming up with the filter to apply after you've somehow managed to present the image that you're trying to get the person to misclassify for 1/20 of a second, followed by a noise mask, is >100 times less valuable at getting them to misclassify the picture than figuring out how to add the time constraint. Not to mention the fact that transformations that actually take cats to dogs and vice versa are necessarily possible, and that differentiation between cats and dogs is an image recognition task with a much clearer limit on human performance than differentiation between cats and laptops, which we are nevertheless amazingly good at.
Cats are very close to dogs, and it seems like this paper has completely forgotten this fact, and is deeply surprised to learn it.
Overall this is a very poor defense of deep learning.
You also don’t need an abundance of murderers.