A study that interviews 135 people? And this is seen as significant? It seems to me saying something like "we tried this on 150 people, and we saw that 60% of them X" proves nothing because of how small the sample size is.
what I am missing?
A study that interviews 135 people? And this is seen as significant? It seems to me saying something like "we tried this on 150 people, and we saw that 60% of them X" proves nothing because of how small the sample size is.
what I am missing?
The sample size you need depends on the strength of the effect in the population. For instance you wouldn't need a sample size of thousands to conclude cyanide capsules were not a safe, effective remedy for tooth ache. Or that people don't like being bullied at work.
This work wasn't statistically focussed, so it's not as crucial, but if you interview 135 people under a good sampling scheme and get very consistent responses about what matters to them it can't be dismissed as meaningless out of hand.
I would expect plenty of room for change in our understanding of meaningfulness as more work is done though...
When you're building an app, a good way to spot problems is to run a user experience test. Sit down at the side of random potential users for half an hour each. Observe how they're using the product, and let them walk you through what they were thinking when you find it interesting.
Past a certain number of users, major UI problems will surface again and again from test to test, alongside marginal - and arguably less important - tidbits of information that let you identify new UI issues. When you reach that stage you know you should stop. It kicks in at around a half dozen users in my experience - and is very rarely over a dozen.
As small as the sample was, I think it was quite relatable.
Also a large amount of studies are done to show a potential area for further research which will involve larger populations when they are approved.
This is where lay and academic understandings of statistics diverge dramatically. There's far more discussion of how to ensure an appropriate sample size (determined generally by the independent variables you're tracking and how you plan on subsetting the sample) and tests for randomness. Which is why the time to call a statistician into your study isn't when you want to run algos over your data but before you even begin collecting data.
Oh, and large sample statistics is a thing. You need to have a minimum number of observations before large-sample methods are generally considered accurate, described by n.
n = 30.
If (and that's a crucial if) your sampling is random, you can draw tremendous inferences from small datasets. This is why many national-level polls rely on samples of about 300 individuals. The key isn't the size, it's making sure those 300 are really, really random. Goof that and you end up with a "Dewey Defeats Truman" headline. Or Trump as the GOP candidate.
https://en.m.wikipedia.org/wiki/Dewey_Defeats_Truman
(Both owe a lot to how phones are used -- in 1948, landline phones were still a sufficient luxury item that it skewed phone-polling methods. In 2016, cord-cutting is having similar effects.)
I had the interesting experience a while ago of looking at data as it came in generating an estimate of a population of 2.2 billion individuals, based on a sample of 50,000 (Google+ activity, relying on a sitemap file of profiles, where sitemaps are restricted to 50,000 entries per). Well within the first 100 records, the long-term trend of ~8-10% (my final value was 9%) of profiles showing any public activity was established. Given web polling and delays introduced (I tried, and succeeded, to avoid tripping any bot-denial mechanisms), the data rolled in slowly, so I simply set my stats script to loop over the incoming files every minute or so.
Somewhere in my testing I also ran a set of resamplings based on the data I'd ingested -- taking small sets of data from within the larger one and looking for any wildly anomolous trends. This would have suggested that the sitemap file itself wasn't randomly constructed, but eyeball tests of various aspects (including account age, region, and activity) strongly suggested it was. This spared me some more complex sample-generation.
(My conclusions were independently verified from a much larger sample of 500k profiles by Stone Temple Consulting and Eric Enge, a pretty heartening validation.)
________
Edit to add: another crucial point is that statistical error is governed by sqrt(n) -- SE = 1/sqrt(n). To halve your error, you've got to square your sample size. So if 30 doesn't work for you, you're looking at 900, not 60. Assuming sampling costs increase with n, reducing error gets expensive fast.
And for the pedants, SE is "standard error", as I'm aware.
Yup.
> To halve your error, you've got to square your sample size.
Nope. You've got to multiply it by 4. (It's the 2 that gets squared.) The square root of 1/120 is half that of 1/30.
Doubling (and repeatedly doubling) occurs through doubled powers of 2. So halving = 2^2, quartering = 2^4, reduce to 1/8, 2^6.
So insteady of n = 30, 60, 120, 240 (initial, halve, quarter, eighth), you end up with 30, 120, 480, 1920.