All the paper shows is (1) stupid ways of counting model complexity and (2) that gradient descent is flawed. Nobody in their right mind believes that increasing the number of hidden neurons can result in a network that is worse on the test set. Since the bigger network contains the smaller network, it is perfectly capable of achieving the same performance, so the only reason why this does not happen is that SGD cannot find it. But of course "SGD cannot always find good solutions" is surprising to nobody, so let's just shit on decades of serious work to get our little paper out.
Sorry for the rant.