The Central Limit Theorem and Its Misuse
lambdaclass.com
lambdaclass.com
It's no use "explaining" what is supposed to be a novel concept by presenting it in a way that is difficult to understand if one is not already familiar with it.
It appears to have been published yesterday-ish. [1]
[1] https://web.archive.org/web/*/https://lambdaclass.com/data_e...
I'm writing this basically because I think it is troubeling how "blogs" work. Explaining, lets say complex stuff, with simple diagrams. Especially in this case it even looks wrong and gives you the wrong intuition.
>Let's consider the following scenario: say we have three coins with different biases (their probability of coming up heads): 0.4, 0.5 and 0.6. We pick one of the three coins at random, toss it 300 times and count the number of heads.
So, not choosing a random coin every time, but choosing it once and tossing it 300 times.
Their property is also called "exchangeable" and it implies identical marginal distributions.
There are too many errors and misconceptions in this blog and thread, confusion between mixture distributions and the sample mean of said mixture, confusion between sufficient and necessary conditions (IID is sufficient to proof the Lindeberg–Levy CLT, but it's not necessary for more general forms Lindeberg-Feller, Lyapunov).
Unfortunately I'm at work and don't have time for a comprehensive rebuttal today. But people with some numpy or R skills can simply do the simulations themselves and see that the blog is not correct.
The procedure given in the example is not a sum of samples from a mixture of Bernoulli distributions. It is a mixture of sums of Bernoulli distributions.
> Let's consider the following scenario: say we have three coins with different biases (their probability of coming up heads): 0.4, 0.5 and 0.6. We pick one of the three coins at random, toss it 300 times and count the number of heads. What is the distribution obtained?
This is approximating the sum of 300 Bernoulli RVs with a Normal, which is perfectly valid.
> As we have seen, if we fix the coin we're tossing, the number of heads can effectively be approximated by a distribution N(300p,300p(1−p)) (where p is the coin's bias). This time, however, each time we take a sample we might be tossing any of the three different coins.
I understand the procedure here to be (1) choosing one of the 3 coins at random and tossing it, (2) repeating step (1) 300 times and summing the resulting number of heads. In this case the CLT does apply: the distribution of the sum-of-number-of-heads is approximately Normal, not the plotted tri-modal density.
This is the procedure. One of the coins is chosen, and that same coin is flipped 300 times. Conditional on which coin was chosen, the number of heads is Binomial, so the unconditional distribution is a mixture of three Binomials.
> This time, however, each time we take a sample we might be tossing any of the three different coins.
which is a different procedure
Yes and no. The sample must be independent and identically distributed. In your case the "identical" part is not correct, as men and women have different distributions (both are normal but with different mean and std). However, if both distributions are normal, then their sum is normal (even with different mean and std).
The fact that the sum is normal in this case has nothing to do with the CLT - it's just a quirk of the normal distribution that the sum is normal. Had men/women had non-normal distributions with different means/stds, then the sum would not be normal.
2) The central limit theorem applies to the distribution of the sample average. It applies whenever the samples are iid and the second moment is finite. The fact that the samples are coming from a mixture of normals doesn't change that.
"Bimodal" doesn't really have a precise definition or test, if you don't assume normal distrubtions.
That paper argues that only if means are separated by 2σ should the distribution be considered bimodal.
But there are many measures of bimodality. [1] [2] [3] [4] [5] [6]
---
In any case, I would be very much surprised if the population couldn't be selected enough (age, race, country, diet, family) to have human height be bimodal by any measure.
[1] https://agupubs.onlinelibrary.wiley.com/doi/abs/10.1029/97WR...
[2] https://journals.sagepub.com/doi/10.4137/CIN.S2846
[3] https://link.springer.com/article/10.1007%2Fs11207-008-9170-...
[4] https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1469-1809....
[5] https://link.springer.com/article/10.3758%2FBF03205709
[6] https://esj-journals.onlinelibrary.wiley.com/doi/abs/10.1007...
For example with stuff that are commonly assumed to be bell-curved like test scores, IQ, etc. What are the iid variables being averaged? Each test question?
in the case of test scores it could be either. for an individual test it's possible the distribution is roughly normal actually. but if you, for example, look at subsets of all SAT scores and take their means (the mean of the subset) then those averages will be Bell shaped because of the clt.
Just to directly answer your question: the iid random variable is the test score or the iq.
The distribution of the original population of iqs in the school is what you're sampling. Each group is a sample. The distribution of the sample mean approaches a normal according to clt as the sample sizes increase.
Now to illustrate the CLT, you roll a die 50 times, and average the result. AND your 300 classmates do the same. If you tally the 301 averages, the distribution of the averages will not be uniform but bell-shaped, with average (approximately) 3.5.
The CLT says (roughly) the distribution of the averages will be approximately normal, regardless of the original distribution.
So it will neither hold strictly for n=50 or n=1, but n=50 may be a better approximation of infinity :)
If you plot the means of the 300 dicethrowers for n=1,2,3,4... dice. You will see the distribution of the means take on more and more of a bell shape. But fully normal it becomes only when n approaches infinity (and then it will be a quite spiked bell of course :))
So, to be clear: the CLT talks about behavior when n approaches infinity. The CLT can be used approximately before that, but it's a very bad approximation when n=1.
So in the original example, we know that if we choose X_n to mean an individual die roll, then the random variable X is not uniformly distributed. However, the CLT tells us that the average of many rolls is indeed normally distributed.
So if you pick n=50, then the average number of pips will resemble a normal distribution. What's the point of having 300 students do it then? It's so we can gather evidence that we're correct :)
Let each individual student do a 50-roll run. We know each run is i.i.d, so let Y represent the distribution of the average number of pips, and Y_1 ... Y_300 represent 300 samples - we can graph the empirical distribution of Y and test it for normality.
However, going back to your actual question, what if we had each person roll just 1 time? Is the outcome normal?
It depends on what you mean - how do we aggregate the 300 individual dice rolls? If you take each individual dice roll and tally it up (how many 1s, how many 2s, etc.), you will find that the distribution is still uniform - so no, a 'sample size' of 1 was not sufficient.
However, what is normally distributed is the average of those 300 dice rolls. However, in this case, we only get a single 'sample' (something close to 3.5), so we lack evidence that the CLT holds (even though it does).
Same goes with your other question - what if you had 50 * 300 people roll the dice? The average number of pips over 15k die rolls is approximately normal, but you've only drawn a single 'sample' in this case.
Going a bit further, this implies that what matters is not just how many dice we throw, but how we choose to define the meaning of those throws - for 15k throws, you can choose to think of it as 15k samples if a single throw, or 15 samples of (the average of) 1000 throws, or anything in between - just pick the definition that's useful.
I was just curious if one can go the other way around, because usually you only have 1 sample of size 15k and not 300 samples of size 50. If you have the raw data it’s just 15k samples of size 1 or 1 sample of size 15k, depending on how you look at it.
Then I could be wrong here, but doesn’t the proof of the CTL also assume that each random sample is the same size? So you can’t have one sample of size 30 and another of size 20 and another of size 25, etc? Each of the 300 samples must be size 50?
Put another way, what you have is 15k dice throws, and the fact that they were thrown by 15k, 5k, or 300 people can be ignored, if you choose to. In fact, it may be useful to 'shuffle' dice rolls into new 'samples' - that gets in to the use of resampling and bootstrap techniques.
And also yes, for CLT to apply, each sample must have the same N.
Smaller samples have higher variance and their means probably won’t converge as fast to a normal distribution.
“Samples” of one, in isolation, have no defined variance and you can’t use them to infer anything.
So what the CLT is saying is that, surprisingly, sums of i.i.d. random variables lose information from the individual variables quite quickly but reliably retain a little data about mean and second moment. Initially a given X_i has all sorts of information associated with it (higher order moments, other distribution characteristics, etc) that disappears as many X_i are summed together. All that is left is information about a mean and variance.
So say I have a situation where a large number of i.i.d. variables are going to be summed together. The CLT tells me that summing N of these variables together is going to be similar to summing together an equivalent number of normally distributed variables (!!). This is because the sum variable is equal in value to n * \bar{x}, but the CLT imposes a distribution on \frac{\bar{x}}{n}.
This justifies why the normal turns up everywhere in practical measurements. A lot of measurements (say, number of people at the beach) are probably really measures of a sum of random variables (maybe there is some non-normal variables that captures the chance a given person goes to the beach). So if the number of people at the beach turns out to be normally distributed (maybe I measure it each day for a few weeks) it isn't shocking. If the total number isn't normally distributed then that implies that there is no i.i.d. variable representing probability a given individual goes to the beach (eg, high correlation between the individual variables).
I didn't quite fail any of the statistics courses I've ever done, YMMV.
[0] https://en.wikipedia.org/wiki/Normal_distribution#Maximum_en...
We have 300 coin throws, so basically a list of H's T's.
My go-to for a histogram would be to have the buckets represent categories. I Throw a coin 300 times and plot the number of heads on one column, and the number of tails on the other column.
Next I would maybe do a bucket-size of 10-throws and count outcomes chronologically in each bucket.
Lastly, I would consider the same, but only with comulative counts.
[1] https://lambdaclass.com/data_etudes/central_limit_theorem_mi...
- Convergence takes infinite time when the sampling variable has infinite variance (imagine a random number generator).
- While the distribution of an aggregate statistic tends to a Gaussian, most underlying distributions are not and most stats is done using the underlying raw data.
Junior data scientists assume everything is Gaussian (sometimes attributing it to CLT) when doing linear regression when most distributions being modeled are not. Then, analysis done using variance on the coefficients is meaningless because the assumptions are incorrect.
To your specific examples - the coefficient of a linear regression is distributed normally (?). Similarly, we know the expected distribution of most maximum likelihood estimators (logistic regression, etc.), and programs will give you the right p-value.
Of course, omitted variable bias is still a problem and it is possible to mis-specify your model. However, I think most data scientists are presenting aggregate statistics (means, regression coefficients) like you said, and that we have a pretty good handle on the underlying distributions.
what does this mean? are you trying to say the code uses the empirical distribution as a proxy for the true distribution?
When does that really even happen in practice though?
> most distributions being modeled are not
That a rather assertive statement. I dare say most distributions are (due to CLT) Gaussian.
I have been doing stats/ML/data-science for more than a decade. The above has rarely been true ever in the datasets that I have looked at. Unless of course the raw data has been transformed in ways to make them look Gaussian. Gaussian is the exception, not the rule.