For example with stuff that are commonly assumed to be bell-curved like test scores, IQ, etc. What are the iid variables being averaged? Each test question?
For example with stuff that are commonly assumed to be bell-curved like test scores, IQ, etc. What are the iid variables being averaged? Each test question?
Now to illustrate the CLT, you roll a die 50 times, and average the result. AND your 300 classmates do the same. If you tally the 301 averages, the distribution of the averages will not be uniform but bell-shaped, with average (approximately) 3.5.
The CLT says (roughly) the distribution of the averages will be approximately normal, regardless of the original distribution.
So it will neither hold strictly for n=50 or n=1, but n=50 may be a better approximation of infinity :)
If you plot the means of the 300 dicethrowers for n=1,2,3,4... dice. You will see the distribution of the means take on more and more of a bell shape. But fully normal it becomes only when n approaches infinity (and then it will be a quite spiked bell of course :))
So, to be clear: the CLT talks about behavior when n approaches infinity. The CLT can be used approximately before that, but it's a very bad approximation when n=1.
So in the original example, we know that if we choose X_n to mean an individual die roll, then the random variable X is not uniformly distributed. However, the CLT tells us that the average of many rolls is indeed normally distributed.
So if you pick n=50, then the average number of pips will resemble a normal distribution. What's the point of having 300 students do it then? It's so we can gather evidence that we're correct :)
Let each individual student do a 50-roll run. We know each run is i.i.d, so let Y represent the distribution of the average number of pips, and Y_1 ... Y_300 represent 300 samples - we can graph the empirical distribution of Y and test it for normality.
However, going back to your actual question, what if we had each person roll just 1 time? Is the outcome normal?
It depends on what you mean - how do we aggregate the 300 individual dice rolls? If you take each individual dice roll and tally it up (how many 1s, how many 2s, etc.), you will find that the distribution is still uniform - so no, a 'sample size' of 1 was not sufficient.
However, what is normally distributed is the average of those 300 dice rolls. However, in this case, we only get a single 'sample' (something close to 3.5), so we lack evidence that the CLT holds (even though it does).
Same goes with your other question - what if you had 50 * 300 people roll the dice? The average number of pips over 15k die rolls is approximately normal, but you've only drawn a single 'sample' in this case.
Going a bit further, this implies that what matters is not just how many dice we throw, but how we choose to define the meaning of those throws - for 15k throws, you can choose to think of it as 15k samples if a single throw, or 15 samples of (the average of) 1000 throws, or anything in between - just pick the definition that's useful.
I was just curious if one can go the other way around, because usually you only have 1 sample of size 15k and not 300 samples of size 50. If you have the raw data it’s just 15k samples of size 1 or 1 sample of size 15k, depending on how you look at it.
Then I could be wrong here, but doesn’t the proof of the CTL also assume that each random sample is the same size? So you can’t have one sample of size 30 and another of size 20 and another of size 25, etc? Each of the 300 samples must be size 50?
Put another way, what you have is 15k dice throws, and the fact that they were thrown by 15k, 5k, or 300 people can be ignored, if you choose to. In fact, it may be useful to 'shuffle' dice rolls into new 'samples' - that gets in to the use of resampling and bootstrap techniques.
And also yes, for CLT to apply, each sample must have the same N.
Smaller samples have higher variance and their means probably won’t converge as fast to a normal distribution.
“Samples” of one, in isolation, have no defined variance and you can’t use them to infer anything.
So what the CLT is saying is that, surprisingly, sums of i.i.d. random variables lose information from the individual variables quite quickly but reliably retain a little data about mean and second moment. Initially a given X_i has all sorts of information associated with it (higher order moments, other distribution characteristics, etc) that disappears as many X_i are summed together. All that is left is information about a mean and variance.
So say I have a situation where a large number of i.i.d. variables are going to be summed together. The CLT tells me that summing N of these variables together is going to be similar to summing together an equivalent number of normally distributed variables (!!). This is because the sum variable is equal in value to n * \bar{x}, but the CLT imposes a distribution on \frac{\bar{x}}{n}.
This justifies why the normal turns up everywhere in practical measurements. A lot of measurements (say, number of people at the beach) are probably really measures of a sum of random variables (maybe there is some non-normal variables that captures the chance a given person goes to the beach). So if the number of people at the beach turns out to be normally distributed (maybe I measure it each day for a few weeks) it isn't shocking. If the total number isn't normally distributed then that implies that there is no i.i.d. variable representing probability a given individual goes to the beach (eg, high correlation between the individual variables).
I didn't quite fail any of the statistics courses I've ever done, YMMV.
[0] https://en.wikipedia.org/wiki/Normal_distribution#Maximum_en...
The distribution of the original population of iqs in the school is what you're sampling. Each group is a sample. The distribution of the sample mean approaches a normal according to clt as the sample sizes increase.
in the case of test scores it could be either. for an individual test it's possible the distribution is roughly normal actually. but if you, for example, look at subsets of all SAT scores and take their means (the mean of the subset) then those averages will be Bell shaped because of the clt.
Just to directly answer your question: the iid random variable is the test score or the iq.