> Some of the problems don’t matter as much if your goal for the model is just prediction, not interpretation of the model and its coefficients.
It depends strikes again.
> As my professor once told our class:
> "If you choose the variables in your model based on the data and then run tests on them, you are Evil; and you will go to Hell."
Again, it depends. If one wants to be pedantic, then choosing to omit any variable in the universe in a non-random fashion is based on some data. Choosing variables via "subject matter expertise" is the same thing but "smarter."
> To explore this I wrote a function (code a little way further down the blog) to simulate data with 15 X correlated variables and 1 y variable.
> mod <- lm(y ~ ., data = d)
If one wants to be pedantic, then why are we using a model that assumes independent variables with correlated data?