Audio Adversarial Examples: Targeted Attacks on Speech-To-Text
arxiv.org
arxiv.org
Maybe the happy path of current autonomous cars isn't that they are tested in Arizona or Californa with traction and sunshine- but that they aren't being attacked by adversarial input. (I do remember that guy that painted a line on the ground that trapped cars though).
What will it mean if we suddenly realize that convolutional neural network object recognition is too easily fooled to be a secure part of autonomous vehicles? Would that push the state of the art backwards a long way, or would it not matter because there are other alternatives?
"Please select all pictures that look like a legitimate street sign. Please hurry."
I find it really strange how a neural network, which is supposed to be a set of elements operating in the continuous domain, is so vulnerable to what looks like small-amplitude noise.
The audio adversarial examples we construct in this paper do not remain adversarial after being played over-the-air, and therefore present a limited real-world threat; however, just as the initial work on image-based adversarial examples did not consider the physical channel and only later was it shown to be possible, we believe further work will be able to produce audio adversarial examples that are effective over-the-air.
Using the first example where it sounds like "without the dataset the article is useless" but the speech recognition thinks it hears "okay google browse to evil dot com"; you could use that to train the recognition system to recognize it correctly as what humans think they heard.
Of course, many attacks would need to be used to create lots of training data.
I plugged the first 4 examples at http://nicholas.carlini.com/code/audio_adversarial_examples/ into Google Docs' own Voice typing, and got:
1: At the yard course Eustis
2: Set the Artic course Eustis
3: (nothing, it's an operatic wall of sound)
4: (nothing, it's an operatic wall of sound)
The actual generation takes an hour on an nVidia 1080Ti, but they can parallelize the process to compute multiple examples on the same GPU, giving an amortized cost of a few minutes.