Another "compromise" is using variable importance outputs from tree-based models (e.g. xgboost), which IMO is a crutch; it says which variables are important, but not whether it's a positive or negative impact.
Another "compromise" is using variable importance outputs from tree-based models (e.g. xgboost), which IMO is a crutch; it says which variables are important, but not whether it's a positive or negative impact.
In my experience this is largely illusory. People think they understand what the model is saying but forget that everything is based on assuming the model is a correct description of reality.
Your typical case of linear regression isn't even close to a correct description of reality, what gets included is largely arbitrary and due to convenience. As a result, different people with different types of data can get very different estimates for any features common to both models.
Also, stuff like this:
"I don't even think simple linear models are actually explainable. They just seem to be. Eg, try this in R:
set.seed(12345)
treatment = c(rep(1, 4), rep(0, 4))
gender1 = rep(c(1, 0), 4)
gender2 = rep(c(0, 1), 4)
result = rnorm(8)
summary(lm(result ~ treatment*gender1))
summary(lm(result ~ treatment*gender2))
Your average user will think coefficient for treatment tells you something like "the effect of the treatment on the result in this population when controlling for gender". I get a treatment effect of 1.17 in the first case, but -0.38 in the second case, just by switching whether male = 0 and female = 1 or vice versa."
https://news.ycombinator.com/item?id=16719754AFAICT, that is also the appropriate role for simpler methods like linear regression, unless you have reason to believe your model is robust to any missing "unknown" features and any included "spurious" ones. In that case you are probably dealing with a model based on domain knowledge and not even running a standard linear regression.
People (over-)interpret the coefficients all the time though, I actually think this is a huge problem.