Geoffrey Hinton: Introduction to Deep Learning, Deep Belief Nets (2012) [video]
youtube.com
youtube.com
Also, batch normalization helps with convergence as well. In addition, LSTMs work when dealing with recurrent neural nets.
I'd stick to the first half to get a good sense of back propagation and working with standard neural nets. I'd hold off on the second half, which delve more into Restricted Boltzmann Machines (RBMs) or autoencoders; these aren't used as much anymore.
To augment your education for things that have happened since 2012, I'd learn about ReLUs rather than sigmoids for activation values, as well as studying up on convolutional neural networks (CNNs) and the recent work in sequence-to-sequence NLP translation via neural networks.