I wonder if one could introduce a secondary classifier which is immune (or more resistant) to this kind of attack as a fail safe. One idea that comes to mind is to back the neural net with a random forest, which I imagine would be very hard to trick with this kind of attack as a collection of independent (key) weak learners are trained on the data. To trick a random forest, you would have to trick the majority of the trees within it.