My experience in CS is that the replicability of experimental results is embarrassingly bad, but this isn't making headlines in the same way so people don't consider it to be a problem.
This is data fraud in psych. Data fraud has also happened in plenty of other fields. When it occurs in these fields it is seen as a one-off. When it occurs in psych, it is because the entire field is useless garbage. That's not the lesson to take away here.
I think that's true for low-tier/low-impact CS papers, but the difference is that literally nobody cares about those papers. The high-tier stuff is easily verifiable (like Tensorflow or whatever) and nobody is writing articles about 5% improvements against a benchmark in some obscure niche scheduling and planning domain.
Outside academic CS people are more scientific about experimental results because it has concrete implications on revenue or spend... but they aren't getting the results from conference papers.
Outside the high-profile cases it seems accepted norm that papers perform far worse when scored against somebody else's benchmark. The real measure of quality is how big the gap is.
Adding to complication in the psych field is the widely observed phenomenon that it draws students with psych problems. Like those who would be willing to fabricate data, for example.
But hey, at least it isn't sociology.
There are resilient exceptions in academia in general, but soft science fields have led the way in the decay of academic standards.
There are dozens of well known fields that are built on epistemological quicksand yet which point-blank refuse to admit to or talk about their problems. Because they have a culture of deny deny deny, even when the evidence is overwhelming, they actually get less attention because why bother doing a nice writeup of fraud if you know the result will just be stonewalling? It's better to focus on fields where there might be some actual response, a chance of improvement, no matter how minor.
Some results do replicate. Big Five personality traits, general classification of mental illness, mainstream IQ results, and human perception and performance. All psychology is not false and there are interesting things to learn from the field.
But the severity and frequency of fraud and non replication is worst in psychology. So much so that half of what you learned in Psych 101 ten years ago does not replicate.
The big challenge for the field is that humans know a lot of psychology. We can track our own thoughts and we are constantly interacting with people and trying to understand their psychology. A biologist who studies ants intensely for 5 years is maybe only one of 100 people have ever watched them that carefully. They'll find lots of new stuff.
Psychology doesn't have powerful techniques than that for determining new truths about humans. It's still mostly give a questionnaire or put people in weird situations.
So there are a lot of researchers hunting around for original ideas that are undiscoverable with our current tech. Some of them are bound to give into the temptation to just make up an interesting result with manufactured data. With p hacking they might not even be committing fraud, they are desperate for a positive result, when the get one they stop looking and publish.
So I would still guess in 2023 that about 80% of "new" discoveries published in journals won't replicate. That rises to 95% with journal publications that are published as university pr and reported in the media.
For coverage of what to remember to unlearn Rolf Degen is great: https://twitter.com/DegenRolf
Popular examples that do not replicate (mostly from Rolf):
Repressed childhood memories Social media harms Search bubbles and echo chambers Women prefer masculine men when fertile Priming (you see a fight in the hallway and then don't cooperate later) Power posing (puff out your chest to feel more confident) Watching eyes make people more honest (in Kahneman and Malcolm Gladwell)
Notice that a lot of these are interesting, and you want them to be true. That's not enough to make it so.
As a bonus, an interesting interview with Daniel Kahneman: https://www.edge.org/adversarial-collaboration-daniel-kahnem...
* Plenty of studies inherently cannot survive 20 years. Send out a survey and ask people "on a scale of 1 to 10 how happy are you". These results will not match the results in 21 years simply because the population you are studying is different. This example is simplified, but explains why many results change over time.
* Plenty of studies have survived 20 years, and those that study psychology are really clear about which knowledge is foundational and which is on more questionable ground. The Big Five personalty traits is about 40 years old and has been found to be consistent across cultures. A ton of psychological research around psychological responses has survived the test of time.
* Psychologist often relay on more questionable data than other scientists because it's much cheaper and easier for psychologist to collect questionable data than other sciences. This allows for a wider exploration and makes research much more accessible. Anyone reading a paper with this type of data will be naturally skeptical. I am specifically referring to data collected from self-reported questionnaires, often collecting a sample that generalizes poorly over interesting populations.
Former semiconductor researcher here. It is quite common for researchers not to believe published papers - even in prestigious journals from well known researchers. No one bothers replicating, and they don't put enough information in the paper to replicate (competition - they don't want others to get the secret sauce).
This was true for both computational and experimental papers.
Oh, and I did know one person personally who falsified data (and was caught). He just transferred to another prestigious school and got his PhD there instead.
It was quite demoralizing and helped me decide not to pursue academia.
But that's not the only problem. The problem is that people try it anyway. They have some hypothesis within a theoretical framework which is embedded in other frameworks, none of which is proven. How could they be? But, they set up an experiment anyway, and (usually after a couple of attempts) they find something that can reject H0. Then they publish an article stating that their theoretical framework is a fact.
There is so much wrong in this process, even ignoring the manipulation of the stimuli and conditions, and the statistical procedures, yet the theory has a good chance to make school and become the subject of research in dozens of psych departments, each contributing articles to the literature about it. After a while, an essentially flawed theory enters the handbooks.
Then, if you want to make a name for yourself as a researcher, and you need publications, you can simply attack older theories. They'll crumble like cookies, you get your publications, and the cycle restarts.
And as critical as I am of psychology, other social sciences have even lower empirical standards. Educational sciences, sociology, linguistics all have very little to show for a century of research. And that's ignoring disciplines like political sciences or history.
Some of it is used in policy. So when some convicted murderer is out on furlough and kills again, am I wrong to wonder if that was some of the top-notch applied psychology at work?
The work here isn't fraudulent because of a mistake or because the author has shaky statistics knowledge, but because it was fabricated.
Some are purely on the popular science side of failure. The 10k of practice was basically thrown out, but the original research still seems good. They never claimed "if you do 10k of work, you will be good."
The "marshmallow test," though, seems completely tossed? Maybe there was something there?
Anchoring and other items? Not sure how well those have survived. :(
https://en.wikipedia.org/wiki/Replication_crisis looks to be a good article, though I haven't finished reading it.
- Stanford Prison Experiment; Not an experiment, abuse was scripted, experimenter constantly intervened, reactions of participants were faked, and there was not even a scientific hypothesis they were testing. Definitely read the paper[0] debunking it.
- Milgram Experiment: (the one where people were ordered to shock actors) No good evidence for it. Researchers did not follow script, implausible levels of agreement between different experiments. Killer line is, “only half of the people who undertook the experiment fully believed it was real and of those, 66% disobeyed the experimenter.”
- Robber's Cave: (the one where two groups of kids immediately formed tribal hatred between one another) The conflict was orchestrated by experimenters and the experiment was actually repeated because the first time the kids absolutely refused to turn on one another. More information at [1].
- At best, weak evidence for implicit bias testing and stereotype threat.
- Weak evidence of "facial feedback" (smiling causes a good mood and frowning causes a bad mood)
- Good evidence against "ego depletion" (the idea that willpower is limited in a muscle-like fashion)
- Mixed evidence for Dunning-Kruger effect
- Questionable evidence for "hungry judge" effect (the idea that judicial sentences are massively more merciful in the morning and after a lunch recess due to "ego depletion" - this is also thoroughly debunked here [2])
- The 10,000 hours of practice leading to expertise idea has been disowned by its proponents
- No good evidence that tailoring teaching to students’ preferred learning styles has any effect on objective measures of attainment.
- No good evidence that brains contain one mind per hemisphere. i.e. the left-brain, right-brain split that people talk about, especially after the link between the hemispheres is severed.
- No good evidence for left/right hemisphere dominance correlating with personality differences.
Per https://danluu.com/dunning-kruger/:
- Most people talking about Dunning-Kruger have no idea what it actually means. The actual purported bias is much weaker than people claim - basically, that everyone either overestimates their ability, but that estimation is still positively correlated with actual ability, or estimated ability has basically no correlation with actual ability and everyone is just guessing.
- Increasing your wealth does in fact make you happier at a predictable rate. There is no "plateau" of wealth or income - what appears to be a plateau is misleading displays of data. In effect, increasing income by a proportional rate will increase reported happiness by a fixed rate. For example, say you make $10, and your happiness is 50. Then, your income increases to $20 and your happiness increases to 60. Then, doubling your happiness is necessary to increase your happiness by 10. It's a logarithmic function; plotted on a standard axis, it looks like a plateau, but plotted on a logarithmic scale and it's a constantly increasing line.
- Hedonic Adaptation (aka the hedonic treadmill) is a myth. Bad life events (divorce, disability, death of a loved one) all have negative long-term effects on happiness. Vice versa for positive events.
[0]: https://www.gwern.net/docs/psychology/2019-letexier.pdf
[1]: https://www.theguardian.com/science/2018/apr/16/a-real-life-...
[2]: http://daniellakens.blogspot.com/2017/07/impossibly-hungry-j...
https://en.wikipedia.org/wiki/Grievance_studies_affair https://en.wikipedia.org/wiki/Sokal_affair