Recommend reading on the topic
- Biggio & Rolia "Wild Patterns" review paper for the thorough security perspective (and historical accuracy cough): https://arxiv.org/pdf/1712.03141.pdf
- Carlini & Wagner attack a.k.a. the gold standard of adversarial machine learning research papers: https://arxiv.org/abs/1608.04644
- Carlini & Wagner speech-to-text attack (attacks can be re-used across multiple domains): https://arxiv.org/pdf/1801.01944.pdf
- Barreno et al "Can Machine Learning Be Secure?" https://people.eecs.berkeley.edu/~tygar/papers/Machine_Learn...
Some videos [0]:
- On Evaluating Adversarial Robustness: https://www.youtube.com/watch?v=-p2il-V-0fk&pp=ygUObmljbGFzI...
- Making and Measuring Progress in Adversarial Machine Learning: https://www.youtube.com/watch?v=jD3L6HiH4ls
Some comments / notes:
> Adversarial attacks > earliest mention of this attack is from [the Goodfellow] paper back in 2013
Bit of a common misconception this. There were existing attacks, especially against linear SVMs etc. Goodfellow did discover it for NNs independently and that helped make the field popular. But security folks had already been doing a bunch of this work anyway. See Biggio/Barreno papers above.
> One of the popular attack as described in this paper is the Fast Gradient Sign Method(FGSM).
It irks me that FGSM is so popular... it's a cheap and nasty attack that does nothing to really test the security of a victim system beyond a quick initial check.
> Gradient based attacks are white-box attacks(you need the model weights, architecture, etc) which rely on gradient signals to work.
Technically, there are "gray box" attacks where you combine a model extraction attack (get some estimated weights) and then do a white box test time evasion attack (adversarial example) using the estimated gradients. See Biggio.
[0]: yes I'm a Carlini fan :shrugs: