A ray of light. MLPs can approximate any nonlinear function in the domain they have been trained on. What is is about the depth that makes DNNs more tractible to train than shallow networks? Is it that the particular tricks that have been developed for DNNs haven't been generalized to work at arbitrary depths? Is it that it is easier for humans to design the abstractions that are used when they are layered? Are you aware of any theoretical work in this direction?