Actually there is no need for gradient decent at all if you have a linear regression problem, because you can find the min/max exactly by inverting your weight-matrix (pseudo-inverse if degenerate). Additionally, second order optimisation, i.e. using the Hessian in addition to the gradient is a well studied problem in the literature. My understanding is that it is in general not worth it because calculating the Hessian is very expensive (see e.g. https://arxiv.org/abs/2002.09018).