Fooling Neural Networks [pdf]
slazebni.cs.illinois.edu
slazebni.cs.illinois.edu
I wonder if enough work is being done to combine the achievements of each field. Whenever I see adversarial examples I wonder why people aren't doing more preprocessing to root out obvious problems with normalization in scale, color, perspective, etc. Also, if we could feed networks with higher level descriptors instead of feeding low-information-density color images, wouldn't that make life easier.
I'm sure I'm not the only one thinking this, is there any good research being done in that space?
This library combines "classic" digital signal processing with a smaller RNN. As a result, it's smaller, faster and probably also has less uncanny edge cases than approaches that use an RNN for the complete processing chain. I think many use cases could benefit from this approach.
- lots of people in the DNN for machine vision community do not have a background in classical techniques.
- a lot of classical techniques and preprocessing pass make no real difference when applied to the input of a DNN and are thus worth eliminating from the pipeline to simplify it (this has been my experience).
However, I do think that there are gain to be gotten by combining classical image processing ideas with neural networks. It just hasn't really happened yet.
The peddlers of such a message were bamboozled by early successes and lacked sufficient experience of empirical science to realise this was never going to work.
No NN will discover the universal law of gravitation from any dataset not collected on the basis of knowing this universal law. With 'statistics as theory' there can never be new theory, as a new theory is a precondition of a new dataset.
Any image on the internet that you save to your phone is now a risk to yourself, even if it looks innocent.
Then it would take a super high level human reviewer to determine that original image, although sexual in nature, is not in fact illegal.
There are so many parties that can submit images to this database, it presumably wouldn't be too hard to subvert one of them with cash or laziness.
One of the big problems with neural networks (and other AI techniques as well) is that they cannot explain their classifications, which makes it difficult to determine whether a classification is correct. Most people seriously underestimate how difficult this task is. Humans can do it quite easily because our hardware has been optimized by eons of evolution. Neural networks are only in their infancy.
In short, it’s not only that you can devise adversarial examples that find the blindspots of the function approximator and fool it into misprediction, it’s that for any learning optimization algorithm you can abuse its priors and biases and create an environment in which it will perform terribly. This is a fundamental and inherent feature of how we go about machine learning — equating it with optimizing functions — and we will need a paradigm shift to go around it.
It’s curious to me how most of these results are known for decades, yet most researchers seem dead set on ignoring them.
This paper provides some interesting results on the weakness inherent in universal priors: https://arxiv.org/abs/1510.04931
Maybe we can work our way backwards from the adversarial examples to the inductive biases?
The interesting tradeoff with ML systems is that you trade lots of individual human crap for one big pile of machine crap. The advantage of the machine crap is that you can actually go in and find systemic problems and work on fixing them at a 'global' level. On the human side, you're always going to be stuck with an unknown array of individual human biases which are incredibly difficult to correct.
If hypercomputation is possible, then anything based on Kolmogorov complexity would be SOL, but if not... is Solomonoff induction just too expensive in practice?
I've made such plots when the input is 2d, breaking the input space into discrete chunks/pixels, having the net classify, and then coloring that pixel according to the classification, and what usually happens is something like what an SVM would produce: large contiguous regions of the same class.
But when the input space is high dimension, and the net is super deep, who is to say what this classification looks like... My guess is it looks less like oil and water carefully poured in a bottle, and more like oil and water shaken vigorously in a bottle.
Do you have any citations about how NNs subdivide the input space, or how regular it is?
The way I have thought of it so far is that we humans subdivide the input space, then stick those blocks into a NN that could have huge Lipschitz bound, and observe the output of a highly irregular function.
When you say "What neural networks do is subdivide the input space and assign a label to it." It sounds more like subdividing the input space helps solve the NNs problem (minimizing the loss). But, it seems to me that that is not so related to minimizing the loss. (Partly because the NN never sees most of the input space during training, and neither is it relevant to what humans want: generalization)
It's not surprising that a network can be fooled by small input changes, but if some image preprocessing is enough to solve this It's not a big problem.
On the other hand, if I can make a sign that looks like a stop sign to people but looks like a road work sign to a tesla, that's obviously a big deal.
These slides touch on the difference by saying that physical examples of adversarial inputs are harder, and they mention some mitigation techniques, but they don't seem to really quantify how effective mitigation is in real world scenarios.
Now we just use n models in production and use voting for produce the label.
As n gets large, does this become robust to adversarial inputs?
Thanks again.
Edit: I guess that's similar to "image quilting" (whatever that is) in this slide deck. This is the first time I've seen something like this mentioned. Seems like a straight-forward solution.
Thanks!
Despite this, recently, we made a 1 minute explainer video introducing adversarial attacks on neural nets as a submission for the Veritasium contest: https://youtu.be/hNuhdf-fL_g Give it a watch!
The only trick which works the most is to revert horizontally the image, at random. When it works, Google is not about to find similar images.
There can't be any universal stimuli simply because there's multiple languages and cultures that don't all have the same response to stimuli.
There's a relationship between neural activity and writing system, for example [0].
Then there's stimuli that activate the language centre in some languages (e.g. click-sounds the Khoisan language families in Africa) but not in others. Some languages (especially East Asian languages like Vietnamese) also use tone to distinguish lexical or grammatical meaning, while Indo-European languages do not. which is another significant difference in (here: verbal) language processing.
All this leads me to conjecture that supernormal stimuli are highly unlikely in this context due to the high-level nature of the subject as well as the differences and the diversity in the involved regions of the brain.
[0] https://www.sciencedirect.com/science/article/abs/pii/S01680...