Discusses some mathematical scaling laws as deep net widths approach infinity, and how it is relevant to tuning large networks. His work shows you can tune hyper parameters on smaller width networks more cheaply and then apply them to larger width ones where it is too expensive to do lots of tuning.