How Adversarial Attacks Work
blog.ycombinator.com
blog.ycombinator.com
Clearly there is a marketing opportunity for t-shirts that make you recognize as other things. Who doesn't want to show up in an image search for toasters on Google Images ? :-)
But my current best guess on how this issue will be addressed will be with classifier diversity an voting systems. While that just moves the problem into a harder and harder to synthesize data set (something that not only is adversarial to one classifier but gives the same answer on several), I believe it will get us to the point where we can trust the level of work to defeat them is sufficiently hard to make it a non-threat.
If the map and road sign disagree, just do what everyone around you is doing (ie, go with the flow) until you recover agreement. It's not perfect, but stretches out attacks to having to compromise long stretches of road and multiple vehicle types. If the cars for the next mile somewhere can find pace, the whole group can; similarly, if any kind of car can find pace, they can all use that to set.
Obviously safety checks and such, but human drivers only loosely follow signs and often get behavioral cues from other cars -- why can't automated cars do the same when they get confused?
It just makes it harder to catch who did it.
https://www.law.umich.edu/special/exoneration/Pages/casedeta...
> The border between “truth” and “false” is almost linear. The first cool thing we can derive from it is that when you follow the gradient, once you find the area where the predicted class changes, you can be fairly confident that the attack is successful. On the other hand, it tells us that the structure the decision function is far simpler that most researchers thought it to be.
Humans can be fooled by optical illusions as well; but those illusions are much more limited and much more noticeable than most of these. My (very non-specialist) interpretation of the italicize clause is that a) vulnerability to these attacks is a continuum, not a binary, and b) the current ease of these attacks reflects the crudity of our current ML techniques.
https://static.independent.co.uk/s3fs-public/styles/article_...
What I find much more disturbing is that even though I measured and known those lines are parallel or that those shades of gray are the same, I can't "unsee" the illusion.
There's also a sense in which this is extremely non-trivial; any agent (human or machine) can be subjected to adversarial attacks. They just look different right now, and our current systems are vulnerable to very simple ones.
It seems to me that improving an algorithm's resistance to adversarial attacks is much more feasible than improving a human's resistance to their own class of adversarial attacks.
DeepXplore: automated whitebox testing of deep learning systems https://blog.acolyer.org/2017/11/01/deepxplore-automated-whi...
There are places in almost every police zone where people driving misinterpret the situation almost perfectly consistently. Which means in some cases that the exact same accident keeps happening with a certain regularity, mostly just involving cars (and generally in that case nothing's done about it), but sometimes involving bikers or pedestrians and in those cases usually after a while the roads are changed so it stops happening.
So here we are:
1) there are "adversarial attacks" on the human mind, in existing road situations
2) those attacks were built into our road network, by humans, and presumably accidentally
3) they systematically lead to accidents and cause deaths, mostly people who weren't even the human misinterpreting the situation. Also, a lot of economic damage is caused.
4) we do one of two things to fix it:
a) we just accept this as a fact of life and don't care
b) we change the road situation until we purely accidentally hit one that doesn't get misinterpreted.
So how to deal with this for AI drivers ... I know ! How about the exact same way ?
AI drivers, I've been stuck behind them in MTV traffic enough to know this, are far safer than good human drivers. They far exceed the ability average human drivers. And they wipe the floor with the very best humans where it comes to patience with fellow road users.
Why do we hold them to a 100% standard ? Nobody and nothing, human or otherwise, matches up to a 100% perfect standard. A stick you use to beat a dog will not have a 100% success rate, once every 10 years or so that stick will break and maybe even injure the guy holding the stick by bouncing around, and yet somehow that is acceptable ...
We need to talk about how good drivers need to be. 100% is simply not an acceptable answer.
[0] https://arxiv.org/abs/1610.05755
[1] https://static.googleusercontent.com/media/research.google.c...
In the ATM example you don't directly load an image of the check to the computer inside the machine. You design a check in photoshop, add the noise, print it out, feed it to the machine, which takes a picture of the check. Mobile bank apps still require you to take a picture of the check so you don't have enough control there either.
Similarly in the road sign example; the lighting, angle between car and sign, dirt on the sign, etc all mean the car sees a much different image than you designed.
I'd think all of these steps mean the classifier gets a dramatically different image than you intend and the attack fails. There's maybe a vanishingly small probability it works when the stars align, but that could be easily mitigated by taking multiple consecutive images and looking for an odd results.
Recent work by me and some friends demonstrates that physical-world adversarial examples can actually be made quite robust, and you can synthesize 3D adversarial objects as well, and make them consistently classify as a desired target class: http://www.labsix.org/physical-objects-that-fool-neural-nets...
From the abstract: "Machine learning (ML) models, e.g., deep neural networks (DNNs), are vulnerable to adversarial examples: malicious inputs modified to yield erroneous model outputs, while appearing unmodified to human observers. Potential attacks include having malicious content like malware identified as legitimate or controlling vehicle behavior. Yet, all existing adversarial example attacks require knowledge of either the model internals or its training data. We introduce the first practical demonstration of an attacker controlling a remotely hosted DNN with no such knowledge. Indeed, the only capability of our black-box adversary is to observe labels given by the DNN to chosen inputs. Our attack strategy consists in training a local model to substitute for the target DNN, using inputs synthetically generated by an adversary and labeled by the target DNN. We use the local substitute to craft adversarial examples, and find that they are misclassified by the targeted DNN. To perform a real-world and properly-blinded evaluation, we attack a DNN hosted by MetaMind, an online deep learning API. We find that their DNN misclassifies 84.24% of the adversarial examples crafted with our substitute. We demonstrate the general applicability of our strategy to many ML techniques by conducting the same attack against models hosted by Amazon and Google, using logistic regression substitutes. They yield adversarial examples misclassified by Amazon and Google at rates of 96.19% and 88.94%. We also find that this black-box attack strategy is capable of evading defense strategies previously found to make adversarial example crafting harder."
Maybe the problem could be mitigated by penalizing "direct" responses to small-scale features while training, which is not a trivial thing to do though. One approach I could think of is training with multiple altered versions of the image, e.g. various amounts of blur, noise and mean/median filters applied. Or the other way around: To be more confident about a result, scale down the image to a fraction of it's size, run detection on that and compare results.
Are such techniques in use, or being researched on? I'm not too much in the loop about those topics.
"Are such techniques in use"
Yes, it's called 'data augmentation': https://medium.com/towards-data-science/deep-learning-3-more...
-- That has to be an overly broad statement (edit: it might true if you say "neural nets" or something specific instead). I would assume they mean any standard deep learning system and maybe any system that is more or less "generalized regression" but that couldn't be "any machine learning system", I mean one could imagine a deterministic model that couldn't be "tricked".
Plus a link would be good (it's not around the quote, I don't know if it's elsewhere in the article).
That is what's being claimed without proof, when they meant "existing neural network systems are known to be insecure"
But (apart from obvious trivialities) I suspect this is a more general problem with interesting properties—the claim the GP is making boils down to something like: a non-trivial[0] machine learning model with large parameter space is, with vanishing probability (as the number of distinct training samples increases), robust to classes of eps-bounded attacks.[1] This seems like a provable claim that is well-defined and (in a weird, handwavy 'intuitive' kind of way) likely to be true... there are just way too many parameters and too much uncertainty in the local minima that we reach when training from such classes of examples to have 'robustness'.
Perhaps I'm totally wrong, though, and someone will come up with a regularizer that prevents all of this from happening, but this seems highly not obvious to me, at first glance.
---
[0] Non-trivial means that, say, it has non-vanishing curvature (Fisher information, assuming the model returns a vector of probabilities) almost everywhere. I hate to be so nitpicky, but I suspect someone would soon comment asking for definitions of all of these things and that the problem isn't 'well-defined.'
[1] Given a (large enough) set of inputs, the max-norm difference between at least one input and the adversarial example must be ≤eps. In other words, if we want the model to misclassify 'turtle' with 'dog', we don't just give a dog picture---the adversarial picture must look like a turtle in a mathematical sense.
(a) brain simulation is fundamentally impossible for some reason (insert your favorite argument for why this should be the case), or
(b) epsilon adversarial examples exist that fool humans (we just haven’t found any).
Personally, I think (b) is more likely than (a), but it seems much more likely that neither is true: that the human visual system is resilient at least to some extent (...for some epsilon), and could hypothetically be simulated by a machine. It would stand to reason that building a resilient system doesn’t require simulating a human, either, or nearly as much computational power as that would take (given that emulation is inherently very inefficient, and that visual processing only takes up a part of the human brain and seems to be done pretty well by animals with simpler brains). But it might require much more computational power than we have today…
On the other hand, if you limit computational power to around today’s level, and consider only potential architectural changes, the statement seems more plausible to me.
Anyways, my final point is in saying that such a claim doesn't seem too far-fetched in any way.
That being said: I could be totally and hilariously wrong.
> On the negative side, we prove that no algorithm can achieve accuracy of ε < 2η in learning any non-trivial class of functions.
Though this is for a very particular type of adversary which has a lot more power than I'm claiming. In my case, though, I'm strengthening the side that the adversary has a very large number (e.g. infinite, as my claim "N->infty") of possible examples in the training set to be eps-close to.
---
Rereading the above: Sorry, I'm not quite sure I understood your point! Are you claiming that there is such a model which is resilient to adversarial examples, and also has a large number of parameters relative to the hypothesis class size? (e.g. say VC dimension?) In that case, my claim would be definitely false, but I have yet to see such a paper (if you have a reference, I'd love to check it out!)
Adversarial examples can be easily crafted for linear regression, SVM, decision trees and k-NN at very least. I can't think about any ML technique that is proven to be secure against them.
There is also an old blogpost from Karpathy describing that the only reason we are not concerned about adversarial examples in linear models is that nobody uses them to classify ImageNet.
I will only accept blanket statements about "all machine learning systems" when there's a mathematical proof.
I don't mean to lessen the importance of the work, just to point out that saying "X fails at Y on subtask Z using the techniques we have today" is very different from saying "All possible X's, current and future, fail at Y."
It is impressive, but that came out a couple of days ago and is not reviewed yet, as far as I can tell.
It’s not obvious from the URL but two of the authors were interns at Google, which could explain why they describe some aspects of InceptionV3 as white-box (because it's not clear if they have re-trained it themselves).
The models in reference are deterministic, most models for classification/regression are deterministic.
Adversarial attacks are showing that the models are chaotic (in the dynamical systems sense), they're very sensitive to their inputs.
edit: Not all models have this issue, it's been shown that all the typical image based convolutional neural networks are susceptible to this issue. My guess is that it's a more general problem of high dimensional inputs.
It also suggests that people concerned about adversarial attacks shouldn't use off-the-shelf pretrained classifiers, where attacks can be trained offline in advance. Similar to hashing algorithms and rainbow tables, maybe a practice of "salting" an off-the-shelf classifier could be effective in dodging attacks.
So by all means, the authors (and a vast majority of researchers) seem to be confident that ML/DL is the road to AGI, hence can "solve" human intelligence (given it is computational)? For how long are we gonna drag the adage that mimicking a human (Turing test) is equal to reaching human levels of intelligence?
I'm not talking empathy or philosophy. How about just folding laundry. Not just one type, not in a controlled environment, but folding any laundry anywhere.
Why does this mean a machine cannot be taught to do it?
> There are a large class of things humans do that they don't understand the mechanics behind,
Sure.
> and for which there also aren't algorithms.
Well, we manage to encode it in our brains.
Not true, we just have different priors and are fooled in different ways. See stage magic, optical illusions, etc...
But we can often recognize them as such, which is important. Actually, on that note, is there work on making machines being able to recognize magic tricks?
Can we? The occult, new age, and religious sections of bookstores suggests otherwise, as do paid horoscope readings, homeopathic medicines, un/lucky numbers (housing and lottery), and shell-game scams.
I think it's also possible for humans to be 100% aware of the "adversarial attack", and still use these types of mediums for light entertainment. This seems to describe many people who occasionally buy lottery tickets for entertainment, and would probably apply to many who attend and produce modern stage magic shows.
(In fact, I notice that some of the top modern crusaders against con artists who do use illusions and paranormal / occult claims for adversarial reasons are stage magicians themselves -- James Randi, Penn and Teller, Derren Brown.)
Yes, it is possible for people to know lottery odds and still play for the excitement. This does not invalidate the claim that many play the lottery with the expectation of winning, nor that people choose numbers superstitiously.
If you add in noise, then you have to train the network to disregard that noise. And the adversarial input will then be features that occur above this noise floor you ignoring.
Adversarial attacks seem to point to overtraining in some sense.
I believe they are linked to the nature of DL (and ML in general) models. We try to capture very tiny manifold of natural images in the space of all possible images. We found a technique that does it well (CNNs). But by the very definition of its training we train it to output some values in a finite number of points. In the same time CNN's output outside of those points are mainly defined by the model's smoothness. We can expect that in a neighborhood of a given image the output will be roughly the same (it allows them to generalize). Far from any point of training set the model can say anything. Adversarial examples basically hint that this smoothness work well only along natural images manifold, once we step outside it is much more chaotic. Or, equivalently, the neighborhood where the CNN gives roughly the same output is a very thin "slice" that closely follows natural images manifold. Why is that? Probably there are some non-trivial topological reasons. But it is exactly what nobody understands now.
I do believe you are not able to control input image bit-by-bit in many real-world scenarios.
It seems like the attack relies on `doping` the input with features present in a different target, or by masking features of the existing target.
Could you so specifically attack a target without knowledge of its features?
If you have the ability to modify the check why not make it actually look like it’s for a 1000000 dollars (ie even to a human).
If you are going to go out and replace speed limit signs to fool self driving cars, it’s probably equally dangerous whether or not the change is obvious, because if it’s way out of bounds a car won’t do it but if it isn’t then it would also fool humans.
Anyone have some more realistic attacks that are specifically possible due to the imperceptible nature of the changes in the input? Most of the harm I’ve heard in examples simply comes from the fact that if you are controlling the input to an ml system, out can get it to do whatever you want even without using any of these techniques: just actually change the class of the input.
Plausible deniability? "I don't know why it took $1 million for a $100 check, look the check clearly says $100 on it, I didn't edit it. Must be a bank glitch."
Humans are harder to fool. But some bank apps allow you to deposit a check by photographing it. Such apps would be fairly easy to attack.
If you limit yourself to making 100$ checks become 1000$ checks and not 1,000,000$ checks, you might even get away with it.
I'm not saying it's not a problem but there are successful defence strategies already in place for many attacks.
So, I just train my new adversary on the "new" model that was trained on the previous adversarial examples. And now we're back to square one.
I suspect the problem of adversarial attacks is a problem of high-dimensional spaces, not of training on particular samples.
Shallow NN's can be fooled just as well, it seems to be more of a problem of linear models in general. Apparently Geoff Hintons Capsule Networks are more robust due to being "less linear" (Ian Goodfellow mentioned this in a recent talk, don't have the references now to back it up)
Right.
If all ML algorithms can be fooled so trivially, this shows the human mind is not an ML algorithm.
Not that I disagree with your conclusion. Just the hypothesis.
not necessarily.
Humans have a lot more knowledge about contextual information that today's ML often misses. For every "image" you classify you use a LOT of side data. That side data has also been acquired through learning.
Even if your actually saw a 1,000,000$ check or a 200 km/h speed limit, you'd know it's probably nonsense. Based on your life experience so far, you know that there's no way that mom just gave you a 1,000,000$ check, or that the city decided to turn that small residential street into a racing track. An image classification ML algorithm doesn't know any of that.
1. All ML algorithms can be fooled trivially. 2. The human mind cannot be fooled trivially. 3. Therefore, the human mind is not an ML algorithm.
But claim number 2 is clearly wrong. Human minds are trivially fooled. Here's one:
http://www.jimonlight.com/wp-content/uploads/2012/02/Paralle...
This is exactly what an optical illusion is.
But the optical illusion example seems very strong because if you say that human perception includes abilities to classify things, there are many examples just like your link that show cases where the classification routinely goes wrong.
It seems possible to me that there are as-yet undiscovered adversarial examples in human perception that we simply don't have any feasible way to search for, and may never be able to construct practical examples of. (It would probably be extremely unnerving to experience one in real life.)
To put this in comp sci language, we can think of DNNs as proofs that certain inputs belong to certain classes. Since no proof system can be both complete and consistent, something outside the proof system can always either provide unprovable examples or adversarial examples.
It remains to be seen if the same holds for artificial learning systems. Of course! The gap between current learning systems and even the brain of a rodent is vast and we have yet no idea how to get there.
I'm not sure what you're getting at here. By definition, optical illusions are perceptual mistakes limited to the optical system.
It is certainly possible for an optical illusion to trick someone into doing something they don't intend to do because they are fooled into believing the illusion. Imagine painting a set of stairs with a disorienting pattern of shading that makes them appear off, leading you to trip and fall down them.
I don't think it's necessarily for a perceptual illusion to hijack your entire belief system in order to be considered an effective adversarial attack.
A true counter example must be similar to making me believe a car is a toaster by manipulating a single cone cell in my eye. Such a possibility is absurd, which leads me to conclude the human mind is not ML.
Pooling layers in CNNs throw away lots of information about spatial co-occurrence of features, which leads to a possibility of adversarial images where adversarial features are scattered all over the place and so they don't significantly affect humans' visual processing.
The conclusion should be "Human mind is not that kind of ML".
The very nature of ML seems incapable of dealing with this dilemma. You either work within a very constrained domain (low VC) to ensure everything is covered, and thus are unable to fit complex datasets, or expand the domain (large VC) and thus overfit the datasets. It's the bias-variance tradeoff, and is insurmountable.
It's really hard to imagine an optical illusion that makes you mistake objects in the physical world for something else- say, panda for a lawn mower or a car for a pigeon, or something like that.
Note that I don't agree that this tells us anything about whether the human brain (or mind) is "like" a machine learning algorithm. To me this question has about as much meaning as asking if the brain is "like" quicksort.
The physical substrates are so clearly different that the only comparison you can make is on the level of capabilities (say, both are Turing-equivalent etc) not that of actual structures. Like, where's the 1's and 0's in the brain?
Sure but people do, for example, mistake each others' faces or voices. You don't need to mistake your friend for a lawnmower for it to be dangerous.
Also, for example, I often mishear my own name when someone else is speaking. Might not happen with every name but it does with some.
So, to be fair, this too is a different thing than what we're talking about.
Here's a physical object that makes you mistake an insect for a plant:
They are only "essential" properties according to the limitations of your particular perceptual system. To a bee that can see UV light, a stick bug may look entirely different from an actual stick. An animal that hunts by scent would find them clearly distinct.
There is no such thing as "true" ambiguity unless the two objects are actually the same thing in all respects. If they are distinct but appear the same, it's because they are overlapping in some respects but not all.
This is a very confused statement. Fist, since we are talking about image recognition we are - by definition - talking about vision in the spectrum that can be captured by a digital camera and encoded in a typical image format. Second, there definitely is such a thing as essential property. It is a matter of correlation with reality, as well as internal consistency. For example, plants are green and have leaves because of the way they use sunlight. So permuting color of all leaves in the picture is fundamentally different from permuting luminosity of some random pixels.
Different things.
What I tried to rebut is your first, weaker assertion:
> The problem with optical illusions like that is that they are, in their vast majority, made of abstract shapes. Most of them play with our perception of distance and depth - and the majority again work on two dimensions, only.
This makes me see a moving 3d object when I am in fact looking at a static flat object; and I can't shake it, even if I try.
Also, from what I can tell, this just happens to work better with a meaningful shape (the image of a T-Rex). It has something to do with how we confuse convex with concave shapes and the T-Rex's snout just happens to be very convenient to demonstrate this confusion (e.g., I think an image of a horse would work as well). I think the illusion would also work with an abstract shape- except it would be harder to find a good abstract shape to demonstrate it as clearly as with the T-Rex.
Nice how this casually demeans peoples' worries. Of course, the (demonstrated) idea that ML algorithm would pick up on, and amplify, discrimination (among other things) isn't "real" to these guys.
Risk-scoring of loan applications actually happens to be one of the few uses of ML in the non-tec sector, and it's incredibly likely that some of them are already denying people's application because they happen to be named "La David" and not "Emil". But maybe the authors just don't consider "ethics" to ever be a "real" problem?
But "change a single pixel and the ATM gives you $1,000,000" is, apparently, a "real" and "pressing" problem.
- Redundant: Anyone who attacks you is by definition your adversary.
- Unhelpful: According to the article, "designing an input in a specific way to get the wrong result from the model is called an adversarial attack." That sounds much closer to spoofing attack ("a situation in which one person or program successfully masquerades as another by falsifying data" -Wikipedia). For example, a turtle masquerades as a gun by spoofing the machine learning system by changing irrelevant visual details.