There's a lot of value in all that. Especially if your deliverable is something that a business is going to use, and not just a Kaggle entry.
There's a lot of value in all that. Especially if your deliverable is something that a business is going to use, and not just a Kaggle entry.
I know it _seems_ that way, but there's a surprising amount of nuance there and I think we're both fooling and limiting ourselves by letting this idea fester.
For one, unlike linear regression, logistic regression estimates aren't collapsible, so you can NOT interpret them as "changing this input by X changes the output by Y". That's only true if your set of covariates is perfect, which is never true, though in practice this interpretation might not be _that_ far off.
Another issue I see is practitioners not being aware of scaled/unscaled estimates; I've seen real papers from AI groups use logistic regression estimates like feature importance rankings, but using estimates in the scale of the original features, and not understanding the distinction when confronted about it.
From a practical sense, I think practitioners are much better served using random forests as their initial exploratory models. Less effort for results that are in practice at least as good as a well-prepped logit. Plenty issues with feature importance there, but not any worse than with logistic regression.
tl;dr: The upshot is that non-collapsibility means that I can't use LR coefficients for things that I don't really need to use them for, anyway. That doesn't feel like a crippling limitation to me.
(Well, also, I have to occasionally pause to cross my fingers and say, "ceteris paribus," under my breath, which does admittedly make people think I'm some sort of weird Harry Potter nut. Which is OK. They're not wrong, they're just right for the wrong reason.)
Nor does it render its coefficients less interpretable than those of most other models. "Less interpretable than OLS" can still be pretty darn interpretable.
I agree with Jake's interpretation of the conditional interpretation of the estimates, but the practical issue is that virtually nobody not well-educated in statistics will do that correctly. In particular, people tend to do exactly what Jake concedes rarely makes any sense, which is comparing estimates across different model specifications.
You and I might interpret these betas just fine, but if we show them to a less stats-y audience, will they?
Feature ranking seems like a clearly safe interpretation of betas, though I've been bitten too often by letting glm (in R) scale my predictors, giving me back estimates on the original scales, and thus incomparable, and seen it happen to others even more. Easy to miss when your original scales aren't all that different.
As far as I can tell, the GP raised no objection to logistic regression, they simply noted that the illustration didn't actually illustrate logistic regression but something else.
Logistic regression is high bias low variance. If you were talking about fairness bias, then resistance to bias comes from logreg being too dumb to recognize complex non-linear patterns. Not necessarily a pro.
With an ANN, your easiest defense against overfitting is to have great big heaping piles of training data. That's something that's hard to come by in many interesting situations.
All the more power to you if a solid simple logreg model (or even no ML at all) is your first deliverable.
They maybe directionally interpretable, but that’s about it.
https://www.stata.com/support/faqs/statistics/marginal-effec...
OTOH: if you assume an iid framework, the probabilistic marginal effects aren't even needed to go from something like "non-bottle blondes have probability p of being haired, bottle blondes have probability q" to "painting the hair of 1000 women will generate 1000*(q-p) jobs on average". Or you can parameterize a Poisson process for rare events and report exponential/Erlang waiting times. And so on.
This isn't a crazy minority position "logistic regression is not interpretable" is truism from basic ML courses, and blog posts all over the internet.
Something is rotten in the kingdom of Denmark.
Logistic regression as not being interpretable was drilled into me by one of the creators of AdaBoost in grad school. As I said, this is widely held position.