Previously, neural networks were trained by taking single steps down the direction of sharpest gradient for the network. However, in deep networks with lots of layers, the backprop algorithm (which I assume you already know about) got stuck in local minima.
Deep learning got started when Hinton observed that a certain way of training restricted Boltzmann machines wouldn't get stuck as easily, and hence by pretraining the network as if it were an RBM and then switching to backprop, it wouldn't get stuck in a local minimum as early.
As I understand it, nowadays the best method looks something like a generalization of Newton's Method, wherein the direction you move takes into account the second differential and not just the first differential or direction of sharpest descent. You move furthest in the directions that curve the least, and move the least in the directions that are most sharply curved. It turns out that this (plus some other tricks) make it way easier to follow continuous gradients in big weird parameter spaces, so now it's possible to train deep nets, which are a kind of continuous gradient in a big weird parameter space.
Tl/dr: People figured out how to move better through the parameter space of neural networks, by taking into account the second differential plus some other tricks. So now we can train deeper nets.