The author mentions he defines over fit as “Test error is always larger than training error”. Is there an algorithm or model where that’s not the case?
So it's a degenerate case, but not of the training set. (And presumably that's partly what they meant by "pedantic".)
However, that's just because I decided the right data to test it onto. So, you can't really say much on a model using that definition.