This is a huge problem and in my opinion is mostly due to bad incentive structures and bad statistical/methodological education. I'm sure there are plenty of cases where there is intentional or at least known malpractice, but I would argue that most bad research is done in good faith.
When I was working on a PhD in biostatistics with a focus on causal inference among other things, I frequently helped out friends in other departments with data analysis. More often than not, people were working with sample sizes that are too small to provide enough power to answer their questions, or questions that simply could not be answered by their study design. (e.g. answering causal questions from observational data*).
In once instance, a friend in an environmental science program had data from an experiment she conducted where she failed to find evidence to support her primary hypothesis. It's nearly impossible to publish null results, and she didn't have funding to collect more data and had to get a paper out of it.
She wound up doing textbook p-hacking; testing a ton of post-hoc hypotheses on subsets of data. I tried to reel things back but I couldn't convince her to not continue because "that's how they do things" in her field. In reality she didn't really have a choice if she wanted to make progress towards her degree. She was a very smart person, and p-hacking is conceptually not hard to understand, but she was incentivized to not understand it or to not look at her research in that way.
* Research in causal inference is mostly about rigorously defining the (untestable) causal assumptions you must make and developing methods to answer causal questions from observational data. Even if an argument can be made that you can make those assumptions in a particular case, there is another layer of modeling assumptions you'll end up making depending on the method you're using. In my experience it's pretty rare that you can really have much confidence that your conclusions about a causal question if you can't run a real experiment.
Do you think CocaCola and the Sacklers had their own unique ideas shared by no one else? That we've filtered all scrupulous people out of industry?
Scruples are an abstraction at that scale.
[0]: Politico "Coca-Cola tried to influence CDC on research and policy, new report states" [https://www.politico.com/story/2019/01/29/coke-obesity-sugar...]
[1]: "Evaluating Coca-Cola’s attempts to influence public health ‘in their own words’: analysis of Coca-Cola emails with public health academics leading the Global Energy Balance Network" https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10200649/
[2]: Forbes: "Emails Reveal How Coca-Cola Shaped The Anti-Obesity Global Energy Balance Network" https://www.forbes.com/sites/nancyhuehnergarth/2015/11/24/em...
... into the multi-billion dollar companies GP is talking about.
That doesn't mean you're a bad scientist, just an unlucky one. But it does mean you can't get tenure.
So it's easy to understand why people fake results to secure a career.
That sounds like a really easy problem to solve. Just treat valid science as important regardless of the results. The results shouldn't matter unless they've been replicated and verified anyway.
We should reward quality work, not simply the number of research papers (since it's easy to churn out trash) or what the results are (because until they are verified they could be faked).
Actually implementing it across the academic world seems much harder.
1. No, there is minimal or no numerical matching between populations of neurons in retina (ganglion cells) and populations of principal neurons in their CNS target (the thalamus). That demolished the plausible/attractive numerical matching hypothesis. I was trying valiantly to support it ;-)
https://pubmed.ncbi.nlm.nih.gov/14657177/
2. No, there is no strong coupling of volumes of different brain regions due to “developmental constraints” in brain growth patterns. https://pubmed.ncbi.nlm.nih.gov/23011133/
That idea just struck me as silly from an evolutionary and comparative perspective. We were happy to call it into doubt.
I suspect many of the comments are being made by damn fine programmers who know right from wrong ;-) a la Dijkstra. But in biology and clinical research, defining right and wrong is an ill-defined problem with lots of barely tangible and invisible confounders.
We should still demand well designed, implemented, and analyzed experimental or observational data sets.
However, that alone is not nearly enough to ensure meaningful and generalizable results. The meta-analyses were supposed to help at this level for clinical trials but have been gamed by bad actors with career objective that don’t consider patient outcomes even a bit.
Highlighting the problem is a huge step forward and it looks like AI may provide some near-future help along with more complete data release requirements.
If you have done biology—- Hot. Wet. Mess. But beautiful.