The underlying problem seems to be that deep neural network classifiers tend to place their classification boundary surfaces very close to data points in at least one dimension in a high-dimensional space. That makes them brittle - perturb the data very slightly in the wrong direction and they move through a boundary into some other classification.
I don't know enough about the subject to know why training does that, or what can be done about it.