After 100 datasets you're virtually necessarily going to come across these misleading conclusions, no?
After 100 datasets you're virtually necessarily going to come across these misleading conclusions, no?
This is even before p-hacking.
Having said that, the study reported on here has significance levels better than 1%:
> The farther away a municipality (in dots) is from a historical mission, the lower its literacy level today. This unconditional relationship is negative and highly significant with a t-statistic of -4.36.
So, notwithstanding the experts here eyeballing the graph and dismissing it as embarrassing, I'd say there is something going on here.
Paper:
Caicedo, F. V. (2018). The Mission: Human Capital Transmission, Economic Persistence, and Culture in South America*. The Quarterly Journal of Economics. doi:10.1093/qje/qjy024
However, I don't really see what this study has to do with multiple testing. As far as I can tell, they had one dataset, and they tested that, and they got a significant result. Nothing fishy there. Sure the data is noisy, but wouldn't you expect that given how coarse the data is, and how many confounding factors there must be?
[1] https://en.wikipedia.org/wiki/Multiple_comparisons_problem
Not saying that happened here! I haven’t even read the paper. But it is a valid concern when mining data for patterns, especially when those patterns have a questionable theoretical basis.
I totally understand you're not saying that's what's happening here, but just to clarify for other readers, I'd say in this case there are a few reasons we don't need quite that level of suspicion:
1. There is a seemingly-sound theoretical basis for the observation (moreover, there doesn't seem to have been any element of fishing. Sometimes researchers think fishing is OK if they just do it once to avoid multiple testing. But then you certainly have multiple testing at the meta level).
2. This study agrees with at least one other completely independent, reputable study [1] (even different methodologies, apparently) that supported the underlying theory.
I do think the problem of multiple testing at the meta level is vastly underrated by the scientific community. Just, in this case, I think there are reasons to reject that hypothesis.
[1] https://scholar.princeton.edu/sites/default/files/lwantche/f...
You're referring to "Type I error" in which a statistically significant result is obtained by chance, not because the effect is real.
This article does not appear to be an example of the hypothetical scenario you proposed. The trend line is not a "horrid-fit," as attested to by the statistical analysis published in the article - which is open access and may be read for free.
This is a scientific article - and unless you're a scientist, it may be difficult to understand the entire article.
The simplest, most conservative form is Bonferroni correction [1], where you just divide the required significance threshold by the number of tests you perform. In your example, if you required significance at the 99% level (1% chance a false positive, or "Type I error"), then for 100 datasets, each individual one would require significance at the 99.99% level (0.01% false positives for each data set). (99.99%)^100 = 99.005%. This does not require any assumptions about the independence of the tests. Independent tests are the worst case, so this over-corrects if there is actual correlation between the measurements.
There are other, more complicated correction methods that do not give up as much statistical power (i.e., produce fewer false negatives), but that may require making more assumptions.
I've heard it called p-value hacking, Wikipedia refers to it as data-dredging.