I seem to recall that there a recent theory paper that got a best paper award, but can't find it.
If I remember correctly, their counter-intuitive result was that big overparameterized models could learn more efficiently, and were less likely to get trapped in poor regions of the optimization space.
[This is also similar to how introducing multimodal training gives an escape hatch to get out of tricky regions.]
So with this hand-wavey argument, it might be the case that two-phase training is needed: A large overcomplete pretraining focused on assimilating all the knowledge, and a second that makes it compact. Other, that there is a hyperparameter that controls overcompleteness vs compactness and you adjust it over training.