> The objective function for linear regression is quadratic (and convex). This explains why Newton's method finds the optimum in one step--the local quadratic approximation turns out to be a global approximation!
> Notice, however, that Newton's method requires us to invert the Hessian matrix (f''), which is expensive in general. I don't think autograd software can get around this limitation.
> The KeepStepping optimizer is performing "gradient descent with line search", which is a more Googleable term.
For more about optimization for solving linear systems (as is the case for linear regression), I recommend Shewchuk 1994, "Conjugate Gradient without the Agonizing Pain" [1], which has some nice geometric insight.
[1]: http://www.cs.cmu.edu/~quake-papers/painless-conjugate-gradi...