Boosting is a clever way of modelling a conditional distribution. The insight behind the success of pre-training is that, for many perceptual tasks, having a good model of the input (rather than the input->output mapping) is key.
I have no delusion that the algorithms that work for training deep networks are anything like what the brain actually does, but I don't care. There are many tasks where deep neural nets are state of the art.
The greedy algorithm bears some resemblance to boosting in its repeated use of the same “weak” learner, but instead of reweighting each data vector to ensure that the next step learns something new, it re- represents it.
I guess that most people however would not think of this interpretation of greedily pretraining deep networks :). (I wonder if mbq had this in mind).
In the same article your point about good models of the input is mentioned, too (only copy&paste a small part of the paragraph):
Unsupervised methods, however, can use very large unlabeled data sets, and each case may be very high-dimensional, thus providing many bits of constraint on a generative model.
The 2006 paper is really an amazing read in my opinion.
Deep learning is not about fixing the residuals of the current chain. Deep learning isn't even about residuals in the first place. It's about (1) finding good representations of your data (aka feature learning), (2) then adding a discriminative model on top and then (3) tuning everything. There is no relation to boosting at all.