Also why do linear regression (OLS) models need gradient descent at all? Cannot you calculate the parameters directly?
y = X β + ε ...and a few assumptions give you... (X^t X)^-1 y = β*
I might be missing something in the blog post.
y = X β + ε ...and a few assumptions give you... (X^t X)^-1 y = β*
I might be missing something in the blog post.
Linear models are much older than computers, dating back to Gauss at least, and they do not have anything to do with gradient descent.
matrix inversion is ~O(n^3)
gradient descent is ~O(np) where p is the number of predictors and n are the observations (n x p matrix).
for lasso, calculating that derivative of the multiplier is not possible (for all points), so coordinated descent is used.