One of the foundational problems with the soft/social sciences is that many of its practitioners have such weak analytical thinking skills (Note below) and the educational requirements for Psychology majors typically just include a few semesters of statistics courses designed to be understandable for them.
Peer reviewers in Psychology generally don't expect authors to go much beyond the basic statistical tricks students are taught in undergrad. From prior discussions with professors in these fields, it's my impression that most are unable to understand just how flimsy such methods are.
Personally, my prediction is that this won't get better until we have more AI-guided analytical analysis software packages that make more robust analysis accessible to those in soft-science fields. This is, something to replace reliance on stuff like R^2 and p- values.
---
Note: Just to provide some numerical evidence for the assertion that Psychology practitioners tend to be weak with their analytical skills, here's [a PDF of GRE scores](https://www.ets.org/s/gre/pdf/gre_guide_table4.pdf). GRE test-takers planning to study "Psychology" had quantitative reasoning (QR) mean of 149. This is pretty bottom-of-the-barrel; even people planning to go for grad school to do "Arts - Performance & Studio" did a little better, with an average of 151.
Andrew Gelman argued that a fundamental problem with peer review is that it is done by peers. Peers whose background and skills are similar to those of the authors, and thus not likely to catch things that were missed.
[When does peer review make no damn sense?](https://andrewgelman.com/2016/02/01/peer-review-make-no-damn...)
I'm sure something could be gained via education. I don't think you necessarily have to be good at math to understand the concepts. But a lot of undergrad courses focus on hypothesis testing and p-values (despite those methods being condemned by the American Statistical Association), and encourage memorization of steps to do simple math, over understanding what any of it means. I think the ad-hoc approach of many machine-learning intros would do better. Maybe programming isn't any easier for a psychology student, but simply hammering the ideas, with short demos of pitfalls, may help.
For "small n" problems like in psychology, psych students are much better off with a statistics background than machine learning. So what I'm advocating is a change in how these classes are taught.
Of course, I've known plenty of students who program via copy and pasting code, and modifying it until it runs as necessary. The equivalent of memorizing steps math steps. So it will take more to solve the problem.
Psychology is a strange science in that it's a mixture of people with very unquantitative backgrounds, and those who deal with very complex math and statistics. What many don't realize is that meta-analysis itself really was developed as a method in psychology (even if it technically has its origins earlier with Pearson). This registered replication work is an extension of that, again being done by psychologists. It's probably safe to say that more empirical and statistical research on the scientific process itself has been done by psychologists (along with statisticians and many public health researchers) than any other discipline.
In any event, replication problems happen in other domains as well. This has been documented empirically. It might be worse in the biomedical domain than, say, physics or chemistry, but it's not limited to psychologists. What I see in the neurosciences per se is just as bad, if not worse (because it's ignored more).
I think ignorance of issues pertaining to overfitting, etc. definitely contributes, but I also think that ignorance is pretty widespread, and the problems can be sort of pernicious in that they don't always operate intuitively.
There's probably different causes at different times. Some of these effects that are the target of replication tests are relatively older, when people were less aware of some of this phenomena. To be fair, some of this stuff is unintuitive: for example, some of the major journals would require internal replications, over several samples, the rationale being that if someone shows an effect in 5 samples with slightly different designs, it's probably "real." It's not like some of these things were just based on single samples (although some of them certainly were). Of course, now people are aware that 5 small samples does not large-sample replication make.
My intuition is that another cause is that academics is flooded with researchers collecting a lot of data, under a lot of pressure to produce positive findings to attract money. Hype, TEDtalks, grants, hypercompetition, and so forth. Academics now is horribly incentivized to be popular and bring in money (from peers, it's important to note), rather than to be correct. Add in a complex subject, like human behavior, and it's a recipe for disaster. Academics is also full of conventions, that attain the strength of power structures; when these are baked in it's hard to change them because you have people who attained power under those conventions in charge (nothing conspiratorial really, just people have their biases and blindspots).
I'm not sure peer review can change a field's direction. Say the top journal in a field used to publish the top 10% of papers. If only 1% of papers meet some higher standard, then if the journal chose to enforce this, they would essentially cease to exist.
And the people who would otherwise have got hired/promoted/tenured based on publications there, will now win these tournaments based on their publications in the second-best journal. Their contests are with others in the same field.
I guess the more hopeful scenario is that, besides the 1 of 10 top papers that are actually solid, there are (somewhere) 9 other solid papers with less-flashy results -- presumably carefully proving things that everyone thinks, not counter-intuitive things which get you a TED talk. Here it's more complicated, as if the journal chose to publish these instead, it seems to me that hiring/promotion/tenure committees may simply start to view it as a less-prestigious venue.