The Magic of Sampling, and Its Limitations
research.swtch.com
research.swtch.com
I bet many would think n=100 would be worthless once the population reaches millions, or especially billions.
One HN-related piece of evidence for that is when I pointed out what margin of error would be for a n=164 survey sample, I got downvoted hard! https://news.ycombinator.com/item?id=8050801
But I saw this hundreds of times talking to customers when I ran a survey sampling product out of YC.
The problem is the sample size relationship with power is at a square, then it involves the effect size and the variance. Quadratic relationships are unintuitive, many order of magnitude differences are unintuitive, and variance is unintuitive. So it’s like a superformula for the human brain to not guess accurately.
I write experimentation software, and with samples in the millions, data scientists still want more power. Then you run some internal experiment with n=20 and it’s like “oh yeah, super significant”.
---
[1] I was listening to a podcast where trolley problems were brought up and the speaker was lamenting how clearly "unethical" and "irrational" your evolved intuition is given that most people will let the train hit 10 men working on the tracks than to divert it and kill 1 innocent. Trolley problems are intellectually interesting for various reasons but jumping to that conclusion is clearly absurd. Your intuitions are shaped by millions of years of genetic and social evolution to precisely be most rational for actual real-life problems. If you were actually standing at that switch you'd be thinking...
* do I actually trust my eyes in this situation? Are the workers on a parallel track and there's no actual problem here?
* if I pull the switch, will it derail the train and kill N+1 people instead of the 10.
* will the workers just notice the train in time and scurry off the track? Or will the train just stop? How good are brakes on a train anyway?
* how much time do judges and juries spend solving trolley problems?
... and while you were paralyzed thinking about these and a million other things, whatever was about to happen would happen and there would be no trolley problem.
Edit: I see that the article describes some of the limitations. I'm curious on how to work with unknown populations. That said, it does have me revisiting some ideas. Looking forward to it.
Create a simulated population with some distribution of a metric & run multiple sampling simulations. You'll be surprised. You can even put in sampling biases as test the impact.
Monte Carlo simulations are a surprisingly powerful tool. I once discovered that FAANG data scientists were mis-understanding statistical significance in a reporting product they made by half an order of magnitude because they didn't understand the impact of observationalmethodology and sampling bias in their product. In my company, we set our own thresholds much larger than what the product recommended.
Of course if your underlying distribution is likely to be Gaussian which is true for many phenomena, you don't need to bother except as a pedagogical exercise.
Allen Downey has a ton of open source books that use this philosophy [0] and Peter Norvig has used Python notebooks in a similar manner (look at the ones in the Probability section) [1].
[0] https://greenteapress.com/wp/ [1] https://github.com/norvig/pytudes#pytudes-index-of-jupyter-i...
You don't have to sample directly. The entire field of Bayesian variational learning exist to deal with that very problem. Look up Markov chain Monte Carlo, Metropolis algorithm, conjugate priors, reparametrization tricks.
That's the magic of random sampling, you don't need to accurately know anything about the underlying population. If you do know things about the underlying population then you can do clever things like stratified sampling to get even more accurate measurements, but that's not necessary. The magic is that a randomly selected group of 100/1000/10000 is unlikely to be too different from the population as a whole, no matter what that population looks like.
You do have to be able to sample randomly--truly randomly[1]--from the population, though, and that's often an issue. Picking 100 people randomly from the population of "likely voters in the next US presidential election" is a very nontrivial thing. To start with, that population is not even very well defined; who is likely to vote changes over time and is difficult to pin down. Pollsters do various things to try to account for this, but if they fail to predict say a surge in young voters their numbers will end up being off.
Even if the population is clearly defined, it's not easy to survey a truly random sample from it. Some people are hard to reach. Some people don't want to talk to you, and whether or not they're willing to talk to you might be correlated with the thing you're interested in (like who they plan to vote for). You can do things to try to correct for that, but again if you get that wrong (and it's very hard to get right) your estimates will be off.
And of course, if you're interested in things that are rare, like third party voters, you need a much larger sample to get an accurate read. If you sample 100 likely voters there's a pretty good chance you won't get a single person who plans to vote for the Libertarian Party candidate.
[1] For the most basic form of random sampling, simple random sampling, you need not just every individual in the population to have the same probability of getting sampled, but every possible sample (i.e. every possible set of 100) needs to have the same probability of being sampled.
Is there something special about exponential functions or is it just my misunderstanding of statistics/calculus at play here for doing this correctly? I assume it’s the latter but I haven’t figured out what I’m doing wrong.
In any case, it sounds like maybe it falls under the "if you're interested in things that are rare" paragraph in my post above. You can always design statistics that are arbitrarily hard to estimate. The things that we're typically interested in estimating in real life, though--averages, proportions, and similar--are typically estimable with reasonable sample sizes.
It doesn't sound like your test statistic is chi-squared distributed, in which case it's not surprising that your samples fail the test, and sampling more just makes the failure more obvious.
> Is there something special about exponential functions
It's not that exponential functions are special; almost any other function would likely also fail the test. Rather, they're insufficiently special. The chi-squared distribution with k degrees of freedom arises from the sum of k independent standard normal-distributed random variables. Some computations (e.g. sample variance of k draws from a normal distribution) can be expressed using such a sum, but others (e.g. sample variance of k draws from an exponential distribution) cannot.
You'll need to switch to a different test statistic and use that test statistic's distribution (which is unlikely to be chi-squared) to compute your confidence intervals.
I’m trying to confirm that if I run this function N times (let’s say 1000), that the frequency of the numbers generated match the expected distribution.
If the sample matches the distribution, by design the p-value is going to be uniformly distributed--i.e. a p-value of 0.01 is equally likely as a p-value of 0.99.
If you know absolutely nothing about your population, the only thing to look at is the mechanism of sampling. Is there some step in the process that would bias selection?
In the real world, you never know nothing about a population (you heard me, Frequentists) and you can check that the known attributes of the population match the sample. If they don’t, that could hint at something wrong.
You don’t even need to know the attributes ahead of time. Let’s say you want to spot check 100 API calls. You could find the ratio of user agents for the whole population and make sure your sample is close (detecting Sample Ratio Mismatch). Same for distribution of response times and so on. Just be aware that the more you look at the more likely you’ll find something weird! You need to correct for that if doing math or keep it in mind if eyeballing it.
That said, I'm curious how it does against different question to the words. For example, is the MOE really the same for such questions as "How many words are more than 5 characters?" and "How many words start with the letter M?" Feels like this should /not/ be the case to me, but I will have fun doing some of the simulations.
That depends on a uniform distribution of the population and an unbiased sampling method. One of the polls in the 2016 US presidential election [1] would shift Trump's position by a full percentage point based on input from a single man depending on which week he participated in the panel.
[1] https://www.nytimes.com/2016/10/13/upshot/how-one-19-year-ol...
If your goal is to be within 10%, that's not a problem.
At one time long ago I thought we can't just rasterize everything, there must be cases where, or as resolutions increase, it won't scale well/make sense. That didn't turn out to be the case. Another story was at one company I was at, they were trying to make raster-driven PCB masks (I think that was how it worked, I was in a different dept). The traces must be continuous and robust so is normally done pen-plotter style. They were able to keep increasing the resolution of the laser rasterizing process to the point where it no longer mattered.
I think now the default view is to work in pixels, voxels, etc. For all we know, the universe could work that way too. Like what's with dark energy? Space being created by a vacuum.