Canvas also seem to have nice histograms in their UI
Canvas also seem to have nice histograms in their UI
Many assume IID or NID to make the problem tractable, but that's not how the world typically works even on the very very large data scale. More things on why many data science teams/groups/fail because too many ppl treat it as a BI/BA organization...I really should step away from the keyboard now. LOL.
Worrying about class imbalance is a classic example of how ignorant many data scientists are about how the statistical properties of the models they're using work.
For example if you are to try to "solve" class imbalance issues in a logistic model by over/under sampling you're actually throwing out important information about the prior probability of an event. Class imbalance in the data, if it reflects the real distribution of observed events, is often valuable information that will give you a better model.
Yet so many data scientists I've known see class imbalance as a major issue and when you ask "why?" it's clear they don't know even the basic principles underlying the models they're using.
There are times we you are forced to deal with a class imbalance issue due to the way this can impact training data size, but these cases are relatively rare in modern data environments.
I can see the argument where ML Engineers and DS w/out a more advanced statistics/STEM background would fail since they would continue down their list of libraries w/in their prescribed toolbox. Granted many problems can be approximated to be good enough, and let's face it, the FAANG/MAANG/whatever companies aren't running things so super critical where a user getting one extra email or ad presentation will cause them serious injury or death.
Btw, appreciate your comments.
For more details I suggest reading Frank Harrell: https://www.fharrell.com/post/classification/
As an example, train a logistic model on imbalanced data with both the real distribution of imbalanced classes and over/down sampled classes and you'll notice that the distribution of the predictions from each model is different.
This is hugely important because if you are using the output of your model as an expectation to feed into another model (for example predict expected value of a new user account given P(purchase) * E[value of purchase]), rebalancing will give you the wrong answer. This can be corrected in a logistic model by tweaking the intercept but doing this requires you to make some assumptions, and it's more artful of a process than one would hope.
You can also see this if you compare the log likelihood of the true data with the given model and the new data. The model trained on the real data will have a notably higher log likelihood.
Then there are separate issues with over/down sampling.
If you are over sampling the standard error on your coefficients will be artificially reduced by the over representation of examples from the one class, leading you to be more sure than you should be about your parameters. The point estimates should be more or less the same, but if you care at all about inference this change in standard error is a big deal. I believe, but would have to double check, that this would likewise impact the effects of regularization on your model.
If you are down sampling you are quite literally throwing out observations, and using less information than you have when modeling. Additionally you'll get the inverse problem which is artificially increased variance in the parameter estimates. Again, even if you don't care about inference you still have the problem that your regularization is very likely not working how you expect.
In my experience nearly all data scientists think they are experts in logistic regression when very few really understand this "simple" model. All of these issues here will manifest themselves in more complex models but understanding how (and how to correct them) is a much trickier question.