I think the reasoning is that optimizing the training performance is "easy", whereas optimizing the test performance is "hard". If you can guarantee that test performance will be close to training performance, then optimizing the test performance becomes "easy".