* Estimators for [Baeysian and Deep learning] approaches usually have to solve intractable optimization problems. Thus, they fall back on approximations and get stuck in local maxima, and you don't really know what you're getting. *
So, with deep nets, how big a problem is getting stuck in local optima?
I mean, my intuition is that it's not magic, you're optimising a system of functions so you'd get stuck in local optima in the same way you get stuck with a single function. Is that generally the case?