Only Train Once: A One-Shot Neural Network Training and Pruning Framework
arxiv.org
arxiv.org
The relative success of attacks on nets to extract their training data support that this happens in practice too.
Generalization performance as it stands now always has to be evaluated empirically.
Basically, evaluate the performance of the network on the validation set, but train it on the training set, and adjust your network structure and hyperparameters accordingly. Networks that "memorize images" will perform poorly.
But yes, you touch upon a very important point, that is the dataset must be sufficiently diverse.
To be fair, humans have this problem as well; if I showed you an octagonal red sign that had the words "GO" inscribed in the middle you may still mistake it for a stop sign at first.