[1] http://www.cs.toronto.edu/~fritz/absps/imagenet.pdf
[2] http://www.image-net.org/challenges/LSVRC/2012/results.html
[1] http://www.cs.toronto.edu/~fritz/absps/imagenet.pdf
[2] http://www.image-net.org/challenges/LSVRC/2012/results.html
"Although I didn't define it in the article, generalization (to me) means that the gap between the training and the test error is small. So for example, a very bad model that has similar training and test errors does not overfit, and hence generalizes, according to the way I use these concepts. It follows that generalization is easy to achieve whenever the capacity of the model (as measured by the number of parameters or its VC-dimension) is limited --- we merely need to use more training cases than the model has parameters / VC dimension. Thus, the difficult part is to get a low training error."