But we also have empirical evidence that they generalize incredible poorly, namely the existence of imperceptible (adversarial) perturbations which can transfer across images and networks and are catastrophically misclassified.
Generalization is a multi-axis scale, not a switch: you can have more or less generalization in many different dimensions. Being terrible at adversarial examples just means that axis is weak.