I think the situation would improve with better teaching of philosophy of science and statistics (this would educate better reviewers too).
I think the situation would improve with better teaching of philosophy of science and statistics (this would educate better reviewers too).
On the other hand, you might expect that new discoveries by nature have less data since data is likely more expensive for brand new research, and by extension a lower likelihood of meeting these sorts of stringent statistical requirements. Decreasing the p-value threshold may be counter-productive if we dismiss legitimate new discoveries due to essentially economic constraints with data gathering, which would have the impact of making it less likely to get funding to pursue the problem in more depth, thereby slowing the advance of discoveries.
I could see the reverse happening, where higher p-value standards lead to normalization of deviance in the form of worse p-hacking.
Statistics is just another method and almost every method can be hacked or abused. Science is not about putting checkmarks in tables but about reading and understanding ideas and reproducing the results. Tweaking some numeric values is not going to help the review process which is fundamentally broken these days.
This is a classic tradeoff between exploration and exploitation in active learning.
If your view of the world is that there are only a very few hypotheses worth exploring, and you have a good lay of the scientific land, then requiring higher bar of proof is probably good.
If it's a new field that's extremely complex and where very little is known of the governing principles, then requiring very high stats could severely slow progress and waste lots of research dollars.
I completely agree that rather than setting arbitrary barriers for significance, it would seem much better to let people actually understand what was found, at whatever significance it was. Even setting up the null model to get a p-value requires tons and assumptions. The better test is reproducibility and predictive models that can be validated or invalidated. That's where the science is, and not in the p.
I am not at all in favor of this proposal, but one thing it may do is stem the tidal wave of misinformation.
The very practical concern is that entire areas of research have been based on studies replicated and backed up entirely through p-hacking and selectively publishing only papers with positive results. This is a proven issue today. See https://en.wikipedia.org/wiki/Replication_crisis for more.
It may be that there is a pendulum that needs to swing a few times to get to a good tradeoff. But it is clear, now, which direction it needs to swing.
It's something that affects a few fields, not all science. And the problem has been completely 100% overblown.
If the problem is that things aren't replicating, changing the p-value cutoff for significance isn't going to fix everything. It can just as easily be a bad null model that's the problem,in which case you can't trust any p-value. The MRI scan problem was closely related to that.
It's a field-specific and null-model specific thing. Broadly changing the a p-value cutoff for everybody isn't going to fix this issue.
Just the fact that so few replications have been published indicates huge cultural problems. When I did biomed, in my tiny area of expertise there had been 1-2 thousand papers published since the 1980s. Out of these maybe 2-3 were close to direct replications. None of those showed the results were reproducible, but no one cared...
Usually there were "minor differences" in the methods so it resulted in stuff like: "Protein P has effect E by acting through receptor R in cell line L from animal A of sex S and age Y when in media M".
However if you changed L, A, S, Y, or M apparently totally different things were going on (there were then supposedly dozens of receptors for each ligand, each receptor having dozens of ligands in different circumstances, etc).
In the end I found that E was nearly perfectly correlated with the molecular weight of P (using data from one of the most cited papers on the topic, in which they specifically claimed there was no correlation with any physical properties of the ligands).
So the effect has nothing to do with specific ligand-receptor interactions at all, but no one cared. Situations like this (with few published direct replications, the ones that are published are contradictory, the results are all being misinterpreted anyway, and everyone just continues on their way when problems are pointed out) are totally standard for biomed. The replication aspect of the issue is really only the tip of the iceberg of problems.
As a Psychology student, this is a well-known initiative: https://cos.io/prereg/
(Though I can't confirm or deny its widespread usage.)
The publication bias is harder, and pre-registration won't solve this. But I think this is a separate issue, and it's important to address each issue in its own right.
I've seen the proposal from TFA before and with my very limited knowledge, I'm still fairly certain it will never come to pass in Psychology, as nearly half of all modern studies have reproducibility issues (!). It would be beneficial to our field, in the way that a band-aid is beneficial to a gaping wound, but it would require a lot more rigor than has been evidently been displayed so far (and more rigor is more work, and time is limited).
So... Don't hold your breath.
(Sorry if my comment sounds pessimistic, I don't know much and I'm open to being corrected. I still have enough critical thought to be skeptical of some researchers' dedication to intellectual rigor.)
This is necessary, but not sufficient. What's needed is a way to know for sure that the hypothesis was not changed after data collection. I think predeclaring the hypothesis is the way to go.
Not that education can fix all these (you can't prevent evil), but if reviewers and journals and conferences started to accept more the negative results, the incentive in lying would quickly decrease. And people would probably start to "disprove" interesting theories, instead of trying to "prove" niche results...
Fabricating data is essentially fraud. And while it does happen, most of the problems with reproducibility are not problems of fraud.
It won't protect against outliers, but removing outliers will not solve most problems. It'll happen, but again, I don't think the majority of irreproducible studies are due to misuse of outliers.
>Plus it's almost impossible to imagine a world where, before any experiment in any field, you predeclare it.
Not at all. I'm not saying you predeclare every experiment - just every experiment you try to publish.
The way it works is:
1. You make observations (i.e. collect data - no predeclaring anything). If you see interesting patterns, you'll form a hypothesis.
2. This is the stage where you predeclare your hypothesis, and the criterion of falsification.
3. You now collect new data and test it against your hypothesis.
The hard part is ensuring people won't use some of the old data and claim they collected after their declaration. It's a hard problem, but not an impossible one.
People are in the habit these days of collecting a lot of data, seeing patterns, and publishing them. That's really not how a lot of early science was done. Once you see the patterns, you need to conduct more experiments to falsify them.
The opposite hypothesis is the null hypothesis which is "gene X is NOT important in disease Y" or "priming DOESN'T affect outcome Z".
Since we assume that most interventions will not affect most outcomes, these are much less surprising and interesting results. They are seen as "water is wet" type of findings and are thus hard to publish because no one is interested.
Now if your hypothesis is something like "X will cause Y to go up" and you actually find it causes Y to go down, that IS publishable. It is only when X has no effect on Y that you will have problems.
This is assuming predeclaring ever becomes the norm.