Machine Learning Notes - Linear Regression
vilkeliskis.com
vilkeliskis.com
I definitely appreciate the simple approach in the article. If the OP is like myself, perhaps he's posting this to better his understanding and leaving artifacts for others to follow as they learn. I have to point out, there's so much more happening in regression. To do it well, read further on it.
As a concrete example of why- The author mentions the R^2 value but doesn't seem to warn that adding more variables to your model will artificially increase it. For this reason, a better value is the "Adjusted R^2" which adjusts for that. Also testing the validity of your model, building it up from scratch, understanding you can't predict outside of the domain of your independent variables, etc.
With that out of the way, I very much enjoyed seeing some of the math behind this. My class was entirely focused on just learning to use a statistical package to run regression. That's perfectly adequate, fine, and all I'll use on a day to day basis. Understanding what's going on beneath the covers has always just enabled me to be more powerful at the given task.
Thanks!
Admittedly you do usually learn it in unsexy statistics classes rather than sexy machine learning classes...
There was a great post on Stats.SE a few years ago about the difference between statistics and machine learning[1]. Leo Breiman once argued that statistics tends to focus more on model fitting and checking, while machine learning looked at prediction accuracy. The exchange between Andy Gelman and Brendan O'Connor is pretty funny. It has been my personal experience however that many people that apply a method that they brand as "machine learning" are not as bothered with assumptions as my fellow conservative statisticians.
But statistics and machine learning are quite similar in foundation. Barring the differences in terminology, as a professional statistician, I find I have as little difficulty read machine learning papers and algorithms as I do reading statistics ones.
http://stats.stackexchange.com/questions/6/the-two-cultures-...
Artificial Intelligence is an umbrella academic term which encapsulates the study and design of intelligent machines. It's not well defined because AI is evolving so rapidly.
Machine Learning is a branch of AI that is concerned specifically with learning from data; the results of learning are usually used to predict future events. (Think linear regressions, random forests, etc.)
Though not specifically a part of AI, Statistics is the field that formed many of the algorithms used in ML. Stats informs ML research design (e.g., how large of a sample size do I need), generates mathematical solutions from proofs and equations, etc. With the rise of big data, it's slowly merging with ML.
Data Mining is a mix of ML, Stats and Data Engineering. It's more concerned with structuring and extracting patterns from data than necessarily learning from it. It is often a task within an ML project.
A finely tuned linear regression is a devastating machine learning algorithm.
However... If linear regression is being used to predict the next value in a dataset, and the caliber of the regression improves as more data is gathered, then it counts as machine learning. For all I know, Netflix could be improving their picks of my movies with a grand regression.
Checking the condition number of the covariance matrix of the LS estimates (X'X)^{-1} is the way R does it.
[1] http://statsmodels.sourceforge.net/stable/gettingstarted.htm...