Machine Learning Done Wrong
ml.posthaven.com
ml.posthaven.com
I can certainly see why this is a good comparison, because it's true that both engineering methods and statistical methods rely on sets of given assumptions, but it's also really important not to take this analogy too far. Engineering is ultimately something that is done in a mechanistic world with primarily deterministic outcomes, whereas statistical modeling is conducted in a stochastic world with probabilistic outcomes, so it wouldn't be good to think about machine learning as predominantly mechanistic in nature (in spite of its name). Of course, a lot of the seven points that follow in the post actually emphasize the importance of stochastic factors (e.g. outliers, variance issues, collinearity, etc), so the author is clearly not making this mistake, but it might be good to clarify for anyone else who is reading.
"6. Use linear model without considering multi-collinear predictors"
This is a great point, and just to expand on it a bit, you can also have situations where you have simultaneity, i.e. two or more of your features or predictors are either functions of each other and/or functions of some third variable. This type of problem is more difficult to detect but can cause serious problems with interpreting the regression coefficients as it's ultimately a type of endogeneity, which means that common approaches like OLS will not be consistent.
Moreover, there's just too much of an emphasis on prediction. Design and analysis of experiments, handling missing data and the context of the data sets, and quantifying one's uncertainty about parameters in a principled manner for robust estimators are very underappreciated skills in the community. Using p values arbitrarily and "95% confidence intervals" based on an unchecked normal approximation is incredibly more harmful than not doing anything at all. There's just so much more to machine learning than supervised learning.
In my experience (all in big data), I rarely seen people use SVM, usual choices are logistic regressions and tree based models. In some finance and insurance industries you are restricted to use only interpretable models, which linear models are.
On the other hand, if you wanna compare logistic regression with SVM. While the detail is pretty tricky. One simplified view is to compare linear SVM which is essentially hinge loss with L2 regularization against logistic regression with L2 regularization which is essentially negative binomial log likelihood loss with L2 regularization. If you plot the loss functions, it's easy to see how they penalize negative & positive cases differently.
"When dealing with small amounts of data, it’s reasonable to try as many algorithms as possible and to pick the best one since the cost of experimentation is low. But as we hit “big data”, it pays off to analyze the data upfront and then design the modeling pipeline (pre-processing, modeling, optimization algorithm, evaluation, productionization) accordingly."
If done correctly, then I agree. But we have to be carefull about overfitting when we try out several models or make an initial analysis to determine which model to use. In this sense, choosing a model is no different from fitting the parameters of the model.
If you aren't doing this correctly, then you can't really interpret the performance of even a single model. Seen people screw this up in so many ways - my favorite recent one that was quite high on HN was someone using the full dataset for variable selection, before doing a training-testing split afterwards.
http://stats.stackexchange.com/questions/9357/why-only-three...
I agree, that non-linear models are often able to beat linear ones, but if you have limited amounts of data feature engineering will always beat clever algorithms.
I have a question on #7. I have not used the co-efficients to mean feature importance but some times get tempted to use them. How do you explain which factors are the most important factors behind some outcome to non-stat people?
In terms of explaining this to non-stats people, you might want to avoid explaining the p-values directly to them (as it's very easy for people to get confused about what p-values actually mean), so instead you might simply show them which features are "statistically significant". In other words, try to explain the results in a qualitative way rather than a strictly quantitative one.
But the article gives two points why you shouldn't use coefficient values to determine feature importance, which I think are only valid to some extent.
>a) changing the scale of the variable changes the absolute value of the coefficient
and
>(b) if features are multi-collinear, coefficients can shift from one feature to others.
Regarding a), well, that's what standardized coefficients are for.
b) is a bit trickier, but most regression models are based on the assumption of non-collinearity. This is of course a problem with real-world data, because you will quite often find some level of collinearity. That's when you (1) test for this issue and (2) look towards multilevel models.
While I'm not an expert of the theory behind sampling, if you do find the need to tweak sampling to align the default loss function and the business metric, I would say doing grid search first, and validate the result with the business insight, e.g. if you find getting the rare cases right is much more important that getting the common cases right, does that align with the business insight?)