The default makes sense.
It's also clearly stated as the very first parameter in the constructor which defaults to L2. The docs also state in BOLD: "Note that regularization is applied by default."
This post is just a case of pebcac.
The default makes sense.
It's also clearly stated as the very first parameter in the constructor which defaults to L2. The docs also state in BOLD: "Note that regularization is applied by default."
This post is just a case of pebcac.
Why's that? Logistic regressions are perfectly reasonable to run for exploratory purposes. There are tons of situations where regularization is not what you want.
> If you don't know to normalise your data prior to that
It's not a question of not knowing. In an unregularized regression, you don't need to standardize your data beforehand. People who knows this but don't realize regularization is being applied by default can easily make this mistake.
EDIT: Apparently this is getting downvoted. To explain a bit more clearly: if you know nothing about the data, then your default should be to run with regularization, since you won't get a useful result if the data turns out to be linearly separable. In the case of linear regression, it's often sensible to start with non-regularized fitting, since that won't mess up your coefficients. But logistic regression is quite simply a different beast altogether, and this is covered in any intro on logistic regression you care to consult.
> Logistic regression without regularization blows up if your data is linearly separable. This is an issue you don't really encounter with linear regression.
Finding that your data is linearly separable is usually quite important - I do see students mess this up though, to be fair.
If you want frequentist inference, non-penalized logistic regression is the way to go.
(disclaimer, I'm predominantly Bayesian, so with proper priors the point is actually moot for me.)
scikit-learn is an ML library. It's explicitly not a stats library. The assumptions and defaults are correct for ML use cases. GP comment is entirely correct.
If you want to do widespread statistics, use StatsModels which is designed for that.
Yeah but logistic regression is from stats. If you're going to implement ClassicalThingFromNearbyField in an ML library, people are right to complain when the definition doesn't match. Not surprising anyone, there's a lot of overlap between stats and ML.
If I implemented a "SimpleFraction" model in sklearn, it would be useful to use "[# this class] / ([# all classes] + epsilon)" using epsilon to regularize, but that would be a crazy name. Or at least it would be crazy if I used a default other than epsilon=0.
The use case you think about is not what sklearn does. You cant even use it, because it returns no model metrics you need for a specified model structure.
No one uses sklearn for stats, and logistic regression here is really not stats, it is a not a ML estimated conditional expectation, it is simply a consequence of a ML loss function. Ane in that sense, regularization is useful.
People expecting to get a stats model out of sklearn will fail immediately, since it does not even return errors, covariances or statistics.
Still, any good data scientist I know actually knows to read the documentation carefully and has learnt regularisation is the default.
For example, look at the neural network classifier model in scikit-learn: https://scikit-learn.org/stable/modules/generated/sklearn.ne... None of these defaults can be assumed. If you showed me a function call with no optional arguments I couldn't tell you, nor anybody else that hasn't read the documentation, what activation is being used, what optimizer, learning rate, etc... In fact when coming across these kind of models during code reviews the nature of these parameters is the first thing I ask about.
To the point of exploratory analysis in some of the parents ,I prefer statsmodels for that purpose. It's not quite up to where the similarly purposed tools are in other languages, but for most of my work where I care about interpretation, it hits the right spot between usability and providing the standard statistical outputs.