If you can’t reasonably get at or use second order information, how else are you going to optimize arbitrary objectives?
Well, come to think of it it, why don’t DL approaches use BFGS instead of gradient descent?
Well, come to think of it it, why don’t DL approaches use BFGS instead of gradient descent?
I think the primary reason that such methods are not used much in practice is memory and computational cost: each function evaluation is expensive and you need to solve a very large system at every iteration.
Also to reply to a sibling comment, you can add momentum and step length adjustments to second-order methods in much the same way as in steepest-descent to help escape saddles. The only difference is how the descent direction is chosen for the optimization.