How Big Data Creates False Confidence
nautil.us
nautil.us
1. It's become easier to do more experiments, so even experts are more likely to produce some bad conclusions.
2. Data has become much more accessible, so people without rigorous stats backgrounds have an easier time abusing the hell out of stats on datasets.
If you can somehow adjust to remove the bias more resolution might still be useful but you can't adjust for data that's not there - suppose the profiler never samples during garbage collection.
Another example is photography where a longer exposure (more photons) can be useful but won't fix fundamental problems with the camera or how the photo was taken. If the camera is out of focus or shakes because you don't have a tripod or image stabilization, a longer exposure won't fix it and might make it worse.
I think I understand what you're trying to say, but I don't agree with how you've phrased it. More data can lower p-values or increase power for a given analysis, but the assumptions that go into using data don't change when you simply have more of it. And those assumptions are everything in statistics.
In fact, I think the temptation of 'more is better' leads to more use of what is easily available which can be highly biased. You also get much more 'significant' results that are more tempting to believe in. It's harder to shy away from a very low p-value, even when you know the sampling may not be appropriate.
I'm a biologist, and as a field we have a lot of issues to address with bioinformatics. Over-eager investigators will pull out interesting tidbits from datasets without considering how problematic that sort of hypothesis generation can be. Fortunately the grant reviewers seem to be well aware of this for the most part. I've heard of many people having grants rejected because they thought they found a 'one in a million' phenomenon, but it turns out they looked a million times to find it :) Good bioinformatics is still firmly grounded in genetics.
I'm confused why you don't agree with my phrasing. To be clear, and from your comment I think you already know this, p-value is not a measure of confidence. P-value is a terrible metric that jumbles confidence with effect size. As you get more data you have the ability to properly detect small effect sizes with confidence; hence, the part of my comment you quoted and seem to disagree with? But people who don't have experience working with large datasets see a low p-value and often think they have a big effect size with reasonable confidence instead of a small effect size with very high confidence. Chalk this up to a "misuse of stats" :)
Statistics are there to answer a specific question, and even then it is going to be wrong when your data is incomplete or you ask a question of your data that it can't answer properly.
Traditional intuition insists this isn't suppose to happen. But that's why statistics isn't physics. Data does not have to be botched or erroneous to get creative. It can all be true, and the backing of your theory may also be valid. The issue is whether your theory itself holds any weight or precedence in light of all other possible theories. So when taking big data and statistics into account, "all possible theories" is the big data picture, not any specific theory. Searching for one theory is already misguided, because information chaos/noise mounts with scale, as other data scientists will consistently tell you.
But if we consider theories as abstractions of evidence, then this should all make intuitive sense. A shitload of theories should emerge from a shitload of evidence.
As your data set grows, unbounded variance grows nonlinearly compared to the valid data. As variance increases, deviations grow larger, and happen more frequently. This causes spurious relationships grow much faster than authentic ones. The noise becomes the signal.
Related: Overfitting: https://en.wikipedia.org/wiki/Overfitting
Overfitting happens when you try add too many variables to your training data. This happens because people think that by adding more data (variables), they can remove bias. What they end up doing, is becoming better at describing the data they have, but not the overall phenomena.
It's counter intuitive but mathematically true.
I think he's correct in discussing it - I find folks propose new features far more frequently than new observations become available.
Trying to explain how they failed to find the twenty needles in the three pieces of straw, they now want to roll forward to a barn-full of bales of hay to try and find less needles!
There is a lot of work to be done to reverse the clinical misunderstanding and misuse of the tools we have at hand, because to be frank I'd say none of us understands them.
Can someone point me to a piece of reading so I can learn about the correct way to do this kind of statistical analysis?