That's actually a great point. Who says this isn't already the case? There isn't much evidence to suggest that things have gotten better after the Why Most Published Research Findings Are False paper in 2005. I think people should be extremely skeptical of anything they read, irregardless of a peer review stamp.
This is a mistake that many HN readers make; they think that if one is equipped with some above-average level of intelligence, one can discern the validity of new research. But this is wrong. It usually takes many years of study before one can begin to clearly understand what is even being said, let alone whether it has any veracity.
People of above-average intelligence often resent this fact, because it suggests that their smart opinion isn't as valuable as the opinion of an expert, but that is simply a sad fact of life. If they haven't put in the work in that field, then they don't know what they're talking about in that field, and they are incapable of applying skepticism in that field correctly. This is precisely and unfortunately why we are forced to place our trust in experts.
As an example of an issue that doesn't require much expertise to spot, I was reading a health paper a couple of days ago that reported on an intervention in a specific population. The paper said words to the effect of "[the intervention] was effective, especially for men". The associated chart not only had (somehow?!) got a legend that didn't match the actual chart lines, but whilst the line for men did indeed go down the line for women was very clearly flat! This should have been reported as "no effect in women but an effect in men" but wasn't. When the wording doesn't match the data being reported, that's a good sign in any field that the researchers know they're treading on thin ice. That particular claim was a correlation/causation fallacy anyway, which is very common in health. They didn't have any proof it was their intervention causing the reduction and there were a bunch of reasons to suspect it wasn't. But the intervention was long term and high effort so it's not a surprise they wanted to find something.
To gauge what sort of p-value suggests a useful model requires familiarity with the area and the broader context of the research.
If a field accepts P=0.05 as significant it means 1 in 20 results can be false positives just by random chance. Now think about how many papers get published, and many of them report more than one thing. The right threshold should really be an order of magnitude lower. It's not OK for scientists to report FPs at that rate, and that's a big part of the reason for declining confidence in science.
That is not the correct interpretation of a p-value. See the ASA's statement on p-values.[0]
>Researchers often wish to turn a p-value into a state- ment about the truth of a null hypothesis, or about the probability that random chance produced the observed data. The p-value is neither. It is a statement about data in relation to a specified hypothetical explanation, and is not a statement about the explanation itself.
Moreover, while it's fine to suspect that p-hacking took place with p-values just under 5 percent, it is merely a suspicion and nothing more. Throwing out p-values you don't like without further evidence is a different sort of violation of the methodology; it is like p-hacking in the other direction.
>The right threshold should really be an order of magnitude lower.
This is false. Again, see the ASA's statement.[0]
>Practices that reduce data analysis or scientific infer- ence to mechanical “bright-line” rules (such as “p < 0.05”) for justifying scientific claims or conclusions can lead to erroneous beliefs and poor decision making. A conclusion does not immediately become “true” on one side of the divide and “false” on the other. Researchers should bring many contextual factors into play to derive scientific inferences, including the design of a study, the quality of the measurements, the external evidence for the phenomenon under study, and the validity of assumptions that underlie the data analysis. Pragmatic considerations often require binary, “yes-no” decisions, but this does not mean that p-values alone can ensure that a decision is correct or incorrect. The widespread use of “statistical significance” (generally interpreted as “p & 0.05”) as a license for making a claim of a scientific finding (or implied truth) leads to considerable distor- tion of the scientific process.
In short, these things often require a strong statistical literacy to interpret, which most people complaining about p-hacking do not possess.
[0]https://amstat.tandfonline.com/doi/pdf/10.1080/00031305.2016...
Yes, obviously the whole notion of a hard threshold is a bit nonsensical to begin with, but as the ASA statement says "Pragmatic considerations often require binary, yes-no decisions". There doesn't seem any way around that. People face decisions like, shall we continue to fund this investigation? Yes/no. Should we recommend lifestyle changes to the public? Yes/no. It doesn't make sense to try and map a P value into a budget, for example. At some point you need a threshold (and likewise for effect size and other things). Of course at some level there is fuzzyness, which is why I said if a paper is mostly reporting 0.049 values then ... and I didn't specify what, exactly because the conclusion should be something like "fuzzily suspicious and should look closer".
So it's fine for the ASA to complain about statistical significance leading to "distortion of the scientific process", but their proposed alternative can be boiled down to doing more peer review, as most of the things they tell researchers to look at are things researchers will never conclude in the negative about their own results, like study design (because if they did they wouldn't have got to the point of calculating a P value to begin with).
Finally, I disagree when you say "That is false". Scientists aren't going to stop using statistical significance thresholds because they do need to make binary decisions at some point, and dropping the threshold they currently use by 10x would immediately yield major improvements in replicability and robustness.
Anyway, put P-hacking to one side if you dislike that discussion, it's fine. It's not actually the thing that bothers me the most when I read papers. A big gap between the prose summaries and what the data [analysis] actually shows is a much more common source of distortion, IMO.
This is precisely the sort of misunderstanding that surrounds p-values, and it is what the statement aims to correct. Lowering the standard p-value threshold does not imply an improvement in these factors. The problems of p-hacking and creating false positives lie in the disclosure of the methodology, not the threshold of p-value. That is the whole point.
This will be my last comment in this chain. I'm just trying to clear up some persistent misconceptions that I see on the internet.
In many cases results can easily be independently verified. This is why it works for AI. If you publish a result, it should come with code anybody can run. If the code doesn't exist or doesn't do what you say it does, you're a fool and everyone can ignore you. If it does, you don't need anyone's stamp of approval to prove it.
But that doesn't work with medical trials or things of that nature where independently verifying the claims is expensive.