Why Most Published Research Findings Are False (2005)
journals.plos.org
journals.plos.org
Also this is called research. You don't know the answer before head. You have limitations in tech and tools you use. You might miss something, didn't have access to more information that could change the outcome. That is why research is a process. Unfortunately common science books talks only about discoveries, results that are considered fact but usually don't do much about the history of how we got there. I would like to suggest a great book called "How experiments end"[1] and enjoy going into details on how scientific conscious is built for many experiments in different fields (mostly physics).
[1] https://press.uchicago.edu/ucp/books/book/chicago/H/bo596942...
I alter Ionnides's conclusion to be instead: "Roughly 50% of papers in quantitative biological sciences contain at least one error serious enough to invalidate the conclusion" and "Roughly 75% of really interesting papers are missing at least one load-bearing method detail that reproducers must figure out on their own" (my own personal observations of the literature are consistent with these rates; I was always flabbergasted at people who just took Figure 3 as correct).
https://grants.nih.gov/grants/guide/notice-files/NOT-OD-24-1...
Unfortunately, not all NIH institutes understand how to evaluate and moderate this key new policy. Oddly enough the peer reviewers do NOT have access to DMS plans as of this year.
The NIH DMS mandates are about the data generated by an award.
You're talking about (almost certainly) fraudsters denying they committed fraud. The vast majority of non-replicable results have nothing to do with these types of errors, purposeful or not.
Fraud requires intent; it's a word that describes what happened, but also the motivations of the people involved. Incompetence doesn't assume any intent at all; it's merely a description of the (lack of) ability of the people involved.
Incompetent people can certainly commit fraud (perhaps to try to cover up their incompetence), but that's by no means required.
> ...insist they just made honest mistakes
If they're lying about that, it's fraud; they're either covering up their unrealized incompetence with fraud, or trying to cover up their intended fraud with protestations of mere incompetence. If they really did make honest mistakes, then it's just garden-variety incompetence. (Or just... mistakes. To me, incompetence is when someone consistently makes mistakes often. One-time or few-time mistakes are just things that happen to people, no matter how good the are at what they do.)
Maybe someone was incompetent but also knew they were cutting corners. Should they get a pass because they claim they didn't mean to do it? We should hold people accountable regardless of intent.
Scientists do not consider a study to be more than an observation. What matters to scientists is the totality of the evidence.
There are several chemistry youtubers who have failed to "reproduce" papers they are working from. Does that mean chemistry is a farce and doesn't work? No, it means some chemist at some point in history failed to write something down, mostly because they didn't even know or realize it mattered. Science is incredibly difficult and we can only hope to be okay at it.
Not sure what rare means in this context. The more important research is, the more likely there is fraud involved. So in terms of size of impact, it's probably very common.
And then if you combine this with poorly done, non repeatable, or inconclusive research being parroted as a discovery...You end up with quite a bit of BS research.
Just because a study doesn't replicate, doesn't make it false. This is especially true in medicine where the potential global subjects/population are very diverse. You can do a small study that suggests further research based on a small sample size, or even a case study. The next study might have a conflicting finding, but that doesn't make the first one false - rather a step in the overall process of gaining new information.
In lieu of actual fraud or a methodological mistake that wasn't represented/caught in peer review, it's still extremely difficult to control for all possible sources of variation. That's especially true as you go further "up the stack" from math -> physics -> chem -> bio -> psych -> social. It is absolutely possible to honestly conduct a very high quality experiment with a real finding, but fail to account for something like "on the way here, 80% of participants encountered a frustrating traffic jam."
Their finding could be true for people who just encountered a traffic jam, and lack of replication would be due to an unsuccessful generalization from what they found.
In what sense do they not? On the assumption that there can be other "worlds" for which math, but not physics, holds?
A geology that isn't anchored to our physical reality seems intrinsically invalid.
But it also doesn’t make it not false. It makes the null hypothesis more likely to be true.
The other is the introduction or loss of critical cofactors or confounders that radically change environment and context.
Think of experiments of certain types before and after COVID-19.
This is a subtle point, but truth or falsity isn't really the issue. The problem with a non-replicable study is that the rest of science can't build on it. You can't build your PhD on top of a handful of studies that turn out to be non-replicable, and so on. It is true you can't build science on outright false statements, either, but true statements that aren't adequately reproducible are also not a solid enough foundation. That may seem counterintuitive, but it comes down to this truth not being binary; even if a study comes to a nominally true conclusion it still matters if it didn't do it via the correct method, or is somehow otherwise deficient in the path it took to get there. Studies are more than just the headline result in the abstract.
But the whole process of science right now is based on building up over time. How could it not be? It has to be, of course. But non-replicable studies mean that the things you're trying to build on them are non-replicable too. It doesn't take all that much before you're just operating in a realm of flights of fancy where you may "feel" like you're on solid ground because of all the Science you're sitting on top of, but it's all just so much air.
However, it is also true that non-replicability is a signal of falsity, and that is simply due to the fact that the vast, vast, vast, exponential majority of all possible hypotheses are false. As is another subtle point, a scientist engaging in science properly should probably not come to that conclusion and may not want to change their priors about something because of a non-reproducible study very much, but externally, from the generalized perspective of "what is true and is not true" where science is merely one particularly useful tool and not the final arbiter, I may be justified in taking non-replicable studies and updating my priors to increase the odds of the hypothesis being false. After all, at the very least, a non-replicable study tends to put an upper bound on the ability of the hypothesis to be true (e.g., if someone studies whether or not substance X kills bacteria Y, and it turns out not to reproduce very well, the lack of reproducibility does fairly strongly establish it can't be that lethal).
This isn't true at all. You could have future experiments set out to prove the opposite, or to dig into what novel and confounding variables are at play causing conflicting results. The fact that you have studies that inconsistently replicate is actually a sign of possible new knowledge if you can find out which undiscovered variable is causing the conflicts. Especially in medicine, you can have a case study that leads to future studies even though that specific case may not be able to be consistently replicated in other n=1 populations.
"e.g., if someone studies whether or not substance X kills bacteria Y, and it turns out not to reproduce very well, the lack of reproducibility does fairly strongly establish it can't be that lethal"
Or there's another factor that hasn't been discovered, such as back in the day before understanding gram positive and negative bacteria was even a thing. If you don't know the subtypes exist, you would see inconsistent results as that is an undiscovered variable that you can't possible account for until it's discovered. Once discovered and controlled for, the substance could be very lethal.
Part of the reason this paper was impactful is that it was short and punchy, took aim at all of medicine rather than a smaller subfield, and didn't require as much mathematical understanding as other papers.
I knew some of the big names that were hit by the replication crisis. And before that I spent some time trying to talk to psychology researchers at a top school about the problems with statistical testing. But they had limited knowledge of stats and didn't want to go out on a limb when everyone else in the field seemed okay with the status quo. A paper like this can be read by everyone and makes a forceful argument.
> It then try to generalize to "research" and avoid this very narrow support to the claim
This is a good point. The methods in medicine and the social sciences are especially weak and prone to these sorts of criticisms. In the physical sciences, often you can run enough iterations of the experiment to overwhelm any prior.
> You have limitations in tech and tools you use. You might miss something, didn't have access to more information that could change the outcome. That is why research is a process.
I totally agree. Science is basically a control system, or a root finding algorithm, or gradient descent. At any time t there is a gap between the best known science and the truth. But the point is that science converges to the truth over time, whereas no other alternative does.
If the foundational assumptions are wrong or impossible to challenge, time t can extend indefinitely. Additionally, it is surely possible that whole sections of science is waylaid and diverges from truth on account of funding, legislation, big personalities, etc.
There's actually quite a big assumption in here -- namely, that the truth is constant over time and throughout space. (This can possibly be weakened slightly, but that's the gist.)
I personally think the laws of physics probably are unchanging, and I certainly hope they are (because that means the scientific method converges). But whether they actually are is not only unknown, but a question that cannot be answered by empirical science.
But from a technical perspective, the math is perfectly capable of describing theoretical universes where the laws of physics change in time or space. For example, where physical constants are smoothly evolving or manifolds that don't look the same in all directions.
The math can tell us what observations we'd expect in those situations, and so far we haven't observed anything to indicate we should relax those assumptions.
If we did observe that the laws of physics were changing, that observation would be considered science in the traditional sense. So it's not that "truth" is shifting, it's that "truth" is a family of equations indexed by some parameter rather than a single equation. That's a similar flavor to the jump from Newtonian physics to relativity.
If you're interested in this topic, a relevant key word is "cosmological principle": https://en.wikipedia.org/wiki/Cosmological_principle
The current models include these assumptions for the same reason they include any other assumptions, because models where these assumptions aren't included don't do any better a job explaining any experimental results. If new experimental data shows these assumptions fail in some cases, then the models can be updated to handle them, same as any other failure has been handled in the past. (Granted each step becomes computationally more complex, so eventually we might hit a limit where it is too complex for a human to understand well enough to operate on.)
I think this is the best way to talk about the difference between science as a process vs science as what people in white lab coats do.
People have a tendency to conclude that we should abandon science if something isn't quite right. It's better to think in terms of what will get us un-stuck from a local minimum/maximum.
This is an important point, but fortunately nature has provided us with a solution. Namely that teenagers are defiant and look for ways to distance themselves from their elders.
This is Planck's principle that science advances one funeral at a time [0]
> A new scientific truth does not triumph by convincing its opponents and making them see the light, but rather because its opponents eventually die and a new generation grows up that is familiar with it ...
> An important scientific innovation rarely makes its way by gradually winning over and converting its opponents: it rarely happens that Saul becomes Paul. What does happen is that its opponents gradually die out, and that the growing generation is familiarized with the ideas from the beginning: another instance of the fact that the future lies with the youth. — Max Planck, Scientific autobiography, 1950, p. 33, 97
The biggest impediments have historically been multi-generational organizations whose power requires limiting access to science. Famously the centralized church in the time of Galileo. More recently the cigarette and oil industries. Things like big personalities and legislative priorities tend to have much shorter time scale and usually allow for science to ratchet forward generationally. Big personalities die, legislators turn over every few years in the best case or in a few generations in the worst case.
Intra-generational power structures are a much larger impediment. Funerals take way too long to happen.
That is strange since to work in science today you have to be extremely compliant for close to a decade. Most people who aren't gets kicked out or removed in selection processes, what you have left is one of the most compliant parts of society.
Disclosure: Old person
Isn't that generally what convergence means? At least in the math sense, we are talking about t at infinity, meaning that in any real world time limited application, there is always a gap.
Everyone claims its different in their field
The journey to discovery is often just as fascinating as the results themselves. The process (the false starts, debates, even the dead ends) can be incredibly instructive and inspiring.
If this is research then it shouldn't be true as it should belong to the 'most' research set.
One simple angle is Ioannidis simply makes up some parameters to show things could be bad. Later empirical work measuring those parameters found Ioannidis off by orders of magnitude.
One example https://arxiv.org/abs/1301.3718
There’s ample other published papers showing other holes in the claims.
https://scholar.google.com/scholar?cites=1568101778041879927...
Google scholar papers citing this
Why is it that microarray true positive p-values follow a beta distribution? Following the citations led to a lot of empirical confirmation but I couldn't find any discussion of why.
More to the point of this rebuttal, though: why would we expect the amalgamation of 70k micro-array experiments' abstract-reported p-values to follow a single beta distribution? And what about modeling the bias-induced bump of barely-significant results?
If there's some theoretical reason why the meta-study can use the beta-uniform model, then I could see this being only a mild underestimation of the proportion of false positives (14%), but otherwise I'm confused how we can interpret this.
I'm personally very skeptical it's that low. What I found during COVID is that entire literatures exist in this sort of weird space where you can't even say if they're true or false because the basic methodologies of the field don't even get you that far. Computational epidemiology never seemed to test its predictions against reality to begin with, so the whole concept of doing P-value analysis on such papers would be meaningless. Based on personal experience I'd feel like Ioannidis is directionally correct.
It's worth noting though that in many research fields, teasing out the correct hypotheses and all affecting factors are difficult. And, sometimes it takes quite a few studies before the right definitions are even found; definitions which are a prerequisite to make a useful hypothesis. Thus, one cannot ignore the usefulness of approximation in scientific experiments, not only to the truth, but to the right questions to ask.
Not saying that all biases are inherent in the study of sciences, but the paper cited seems to take it for granted that a lot of science is still groping around in the dark, and to expect well-defined studies every time is simply unreasonable.
Why most published research findings are false (2005) - https://news.ycombinator.com/item?id=37520930 - Sept 2023 (2 comments)
Why most published research findings are false (2005) - https://news.ycombinator.com/item?id=33265439 - Oct 2022 (80 comments)
Why Most Published Research Findings Are False (2005) - https://news.ycombinator.com/item?id=18106679 - Sept 2018 (40 comments)
Why Most Published Research Findings Are False - https://news.ycombinator.com/item?id=8340405 - Sept 2014 (2 comments)
Why Most Published Research Findings Are False - https://news.ycombinator.com/item?id=1825007 - Oct 2010 (40 comments)
Why Most Published Research Findings Are False (2005) - https://news.ycombinator.com/item?id=833879 - Sept 2009 (2 comments)
There is an error in Table 2. A set of parentheses is missing in the equation for Research Finding = Yes and True Relationship = No. Please see the correct Table 2 here.When you spent an entire week working on a test or experiment that you know should work, at least if you give it enough time, but it isn’t for whatever reason, it can be extremely tempting to invent the numbers that you think it should be, especially if your employer is pressuring you for a result. Now, obviously, reason we run these tests is precisely because we don’t actually know what the results will be, but that’s sometimes more obvious in hindsight.
Obviously it’s wrong, and I haven’t done it, but I would be lying if I said that the thought hadn’t crossed my mind.
I thought the whole point of doing experiments was to challenge what we "know" so we can refine our understanding?
If the claim is false, though, you can still sometimes get research to support it. If you or the researcher stands to profit from the false claim, then there is a conflict of interest.
In reality, scientists are highly motivated (i.e. biased) individuals like anyone else. Therefore science cannot be done effectively by individuals.
The system that derives truth from experiments - the actual scientific system - is the competitive dynamic between scientists who are trying to tarnish each others' legacies and bolster their own. The scientific method etc. primarily makes scientific claims scrutinizable in detail, but without scrutiny they are still highly liable to produce false information.
Personally, I think the “highly” in your statement is quite over exaggerated. Humans can be convinced to produce bad science, for sure, and there are even journals set up by religious orgs that specifically exist to do just that.
But at the same time, science landed humans on the moon.
Except that the entire point of the article here is that it's not exaggerated.
> But at the same time, science landed humans on the moon.
Cherry-picking a highly successful, well-known example doesn't prove a point.
There must be hundreds, if not thousands, of successful scientific discoveries that went into something as complicated as the moon landing, and if you still don't think that's convincing, just look at the world around you - which looks just radically different from the world of, say, just a couple hundred years ago.
As if our lifespans and quality of life haven't been drastically improved by modern medicine.
I mean, we can cut people open and replace entire parts of them and they're fine. They don't even get sick anymore - thanks germ theory and aseptic technique! Do you not understand how much of a marvel that is?
Before that, people used to get cuts and scratches and just... die. We can now fully rummage inside an arbitrary person's internal organs.
And don't even get me started on long-term illnesses. High blood pressure and cholesterol has been killing humans since forever, and we have medicine that just fixes that. And now, we're getting medicine to rewire our brains to prevent addiction in the first place (semaglutide)
Tell me of a better method to get to the truth. Go on.
That said, wonder drugs are few and far between. The GLPs are at least a once-in-a-decade breakthrough, so that’s probably most of the noise you’re hearing (there are a lot of brand names already).
No one is under the illusion it’s perfect or ungameable. A drug slipping by every few years is bad and often tragic, but IMO nowhere close to indicative of a systematic problem. It is a system that is worthy of a high degree of trust.
Shouldn't we expect some small percentage of failures in these processes given that they are driven by statistics and confidence intervals? Is that even a failure of the process, or is it a known limitation given how much resources and time we are willing to allocate to the discovery process?
As patio11 says, the correct amount of fraud in a financial system is not zero, and the correct amount of false positives in drug approvals is not zero.
Yes, and this is really solved at a more local level. Doctors aren't prescribing new drugs like candy. They, too, are skeptical of their success and will reserve those prescriptions for the most desperate cases. Over years, we (and the doctors) learn how effective these drugs are and what potential side effects they have.
Cardiovascular safety was tested in the original trial. It passed. Nothing in the data during development suggested it was an issue. But trials can’t detect everything.
It wasn’t until it got to market did a safety signal pop up. Then retrospective analyses of large data sets proved it.
Republicans and Russian bots WANT you to hate science and academia and they have frequent pushes across social media platforms to make sure you do.
That was engineering. Closely linked to science, but not the same process of inquiry.
It's not. Markets are good at that. They're actually competitive. Academia is good at producing enormous volumes of documents that claim to be information, and may or may not be if you test them. It's "competitive" in a weird way where people don't compete over what's actually true but over who can convince central planning committees to give them money, which is very different.
You can't separate the production of correct information from the production of goods and services. They're inherently intertwined, the attempt to separate them is how we ended up with such a polluted scientific literature. Having formed a hypothesis as to what is true, you have to test that out in a robust way where you can't easily cheat and you can't easily cheat yourself, even when "yourself" refers to the vast institutional structures that employ you. In other words you need a system that stops you cheating even if that's what would please your boss, your vice chancellor and ultimately the President.
We only seem to have one such system and that's markets. If you cheat then the product or service you provide will be based on beliefs that are false and - eventually - customers will abandon you because the thing you're selling doesn't solve their problem. This is effectively a referendum of the customers. There is no such feedback loop in academia. There are attempts to approximate or emulate it with things like peer review, but they're all shadows of the real thing.
Meanwhile you can encourage cooperation even between competing entities in lots of ways, and it often emerges naturally even in the absence of any specific social policy. Open source collaboration is one obvious example in the tech sector, patents are a more formalized system.
> The system that derives truth from experiments - the actual scientific system...
Yes!
> ... is the competitive dynamic between scientists who are trying to tarnish each others' legacies and bolster their own.
Hm. To some degree, sure, that is one dynamic, but (a) this leads to/presupposes a truckload of perverse incentives and (b) this is not inherent in the system if we rearrange incentives
Of course the implementation is far from perfect. For example, the interaction between impact factor and grant funding produces pressure toward ideological conformity and excessive analytical “creativity”. But the underlying principle of competitive scrutiny is probably a desirable one.
Cooperation is also an extremely fit behavior in natural selection.
If someone has a large bag of money laying around the plan is this:
There are lots of companies that will run material A though machine B for you. There are a lot of science machines. One is to put a lot of them into a large building and make a web page where one can order the processing of substances in a kind of design your own rube goldberg machine.
It can start with all purchasable liquids and gasses, mixing, drying, heating, freezing, distilling etc and measure color, weight, volume, viscosity, nuclear resonance etc, microscope video, etc. Have as much automation as possible, collect all the machines. A robot cocktail bar basically.
Work your way up to assembling special contraptions all ordered though the gui.
Jim can have x samples of his special cement mixture mixed and strength tested. Jack can have his cold fusion cells assembled. Stanley can have his water powered combustion engine. Howard can have his motor powered by magnets. Veljko can have his gravity powered engine. Thomas can have his electrogravitics. Wilhelm can have his orgone energy.
or not... hah....
If any people are involved they should not know what they are working on.
It wont be cheap but then you get an url with your nice little test report and opinions be damned.
But if you want to without human error/bias there is nothing close to removing all the humans.
Things that are controversial, unbelievable or unlikely may have big implications and risking your career on it is usually not a good idea - for you.
Though automation one might drive the prices down to make the brute force approach viable but with somewhat intelligent machines one could also make educated guesses in volume.
You could auto suggest similar experiments while the researcher types their queries complete with prices.
The original question was: How can we do more research without increasing the number of scientists.
I’ve never fabricated numbers for anything I’ve done, but there certainly have been times where I thought about it, usually after the fourth or fifth broken multi-hour test, especially if the test breakage doesn’t directly contradict the hypothesis.
Although, contrary to what I was taught in elementary school, most of the experiments in the physics department of my university didn't even really have a hypothesis. They were usually either of the form "we are going to do this thing, and see what happens", or "we're going to measure this thing more accurately than anyone before".
An example has been times where I really want to use a certain concurrency style, and I'm convinced that it should be faster than the way we're doing things before, so I will write a few non-trivial tests to make sure that's right, and I'll get inconclusive numbers.
Of course, that is still very frustrating if you are on a tight deadline and all you have is one thing you know doesn't work.
Plus these days there's a lot of pressure to run universities more like businesses. To eat, academics have to hit certain numbers, so you see behaviors common in business like faking the KPIs.
A significant result in an experiment, according to Fisher, is just an experience to add to the mental pros-and-cons list. It is not defintive proof of anything.
The startling rise in the publication of sham science papers has its roots in China, where young doctors and scientists seeking promotion were required to have published scientific papers. Shadow organisations – known as “paper mills” – began to supply fabricated work for publication in journals there. https://www.theguardian.com/science/2024/feb/03/the-situatio...
The number of retractions issued for research articles in 2023 has passed 10,000 — smashing annual records — as publishers struggle to clean up a slew of sham papers and peer-review fraud. Among large research-producing nations, Saudi Arabia, Pakistan, Russia and China have the highest retraction rates over the past two decades, a Nature analysis has found. https://www.nature.com/articles/d41586-023-03974-8
That's why a recent article https://news.ycombinator.com/item?id=41607430, where the measurement of China leads world in 57 of 64 critical technologies was based on number of journal citations, was laughable.
Of course the same thing is happening in the 'Western' world too, with a publication ratchet going on. New hire has 50 papers out? OK! The next pool of potential hires has 50, 55, 52 papers out, so obviously you take the 55 papers-person. You want outstanding people! Then the next hire needs 60 papers. And so on.
Paper mills are bad but mostly from the perspective of academic institutions trying to verify people's credentials/resumes. Paper mills aren't really that much of a concern in the sense of published research results being false in the way the article is talking about because people aren't really reading the papers they publish. In that sense it doesn't really matter if there are places where non-scientists need to get one paper published to check some box to get a promotion, because nobody is really considering those papers part of established scientific knowledge.
On the other hand, scientists intentionally (by actually falsifying data) or unintentionally (as a result of statistical effects of what is researched and what is published) publishing bogus results in journals that are considered legitimate which aren't paper mills actually causes real harm as a result of people believing the bogus results, and unfortunately the pressures that cause that (publishing papers quickly, getting publishable results, etc.) exist everywhere, and definitely not just in China, nor did they originate in China.
Let me give just one example of how prevalent the culture of fakeness has pervaded through China. Nowadays, because the economic decline, people are eating out less, and restaurants are getting less and less traffic. Therefore, they needed to cut costs. So some restaurants started using pre-packaged food, and just heat those up in the microwave and serve them up as cooked dishes. Because other restaurants couldn't survive without doing the same cost-cutting behavior, they've all started doing the same things. Thus, most restaurants in China are now serving pre-packaged food. And there's a backlash from consumers, so now even less people eat out. And then restaurants started using expired pre-packaged food. Oh, and because expired pre-packaged food has a tendency to cause diarrhea, some restaurants in China have started adding Loperamide into the dishes to prevent diarrhea.
Fake it until you make it out of China mentality.
Foreigner caught a Chinese couple scooping up gutter oil https://www.reddit.com/r/interestingasfuck/comments/1eo2wmy/...
Gutter oil used to be a major issue in China but the Chinese government cracked down on it a lot a few years ago.
I recommend watching this video about it: https://www.youtube.com/watch?v=G43wJ7YyWzM
I do believe that there exists an insane amount of (STEM) questions where there exist very good reasons to do research on - much, much more than is currently done.
---
And by the way:
> This is what happens when Silicon Valley execs, trying to make their employees more replaceable, call for more STEM education
More STEM education does not make the employees more replaceable. The reason why the Silicon Valley execs call for more STEM education is rather that
- they want to save money training the employees,
- they want to save money doing research (let rather the taxpayer pay for the research).
So what you're saying is that they push for STEM education to make their employees more replaceable...?
A general rule of thumb is rather that better education and/or specialized knowledge makes employees nore productive, but also less replaceable.
Only when they are the only ones that have that knowledge, not when teaching it becomes rote.
> - they want to save money training the employees,
> - they want to save money training the employees,
> - they want to save money doing research (let rather the taxpayer pay for the research).
means they want to offload costs to the public in order to increase profits, which is what I said above.
Maybe I’m missing something, but I do not believe that is the way it is stopped to go. Btw, she has a PhD and failed up into a global scale.
I’ve been meaning to find out if there are any open tools to evaluate someone’s dissertation.
It was equal part stunning and seemingly a bit traumatizing to me considering I still remember it as if it had happened earlier today. I think what surprised me too was her open admission of it, even with external parties present.
No doubt that in the case of physics and chemistry and the like, testing can be a lot longer.
Does the world really want/need such a system? (The answer seems obvious to me, but not above question.) If so, how could it be designed? What incentives would it need? What conflicting interests would need to be disincentivized?
I think it's been pretty evident for a long time that the "peer-reviewed publications system" doesn't produce the results people think it should. I just don't hear anybody really thinking through the systems involved to try to invent one that would.
https://pubpeer.com/publications/14B6D332F814462D2673B6E9EF9...
So while "if it's any good, people will use it" is true and quality contributions will be useful, the converse is not true: the use or reach of published work may be only tenuously connected to whether it's good.
Reputation signals like credentials and authority have their limits/noise, but bring some extra signal to the situation.
I admit to missing the joke in the first reading.
pseudonyms may prevent the abuse of invalid papers by removing the ability of the authors to front institutional reputations for partisan claims.
the movement for science and data to drive policy outside of their domains sounds nice until you find that the science and data are irrepreducible, and the institutions have become laundering vehicles for debased opinions that wash the hands of policymakers. as though the potential for abuse has become the value.
maybe it's a rarefied kind of funny, but the kernel of truth it reveals is that it could be time to start using pseudonyms in some disciplines to make the axis of policymakers and academics more honest.
"If your experiment needs statistics, you ought to have done a better experiment" - Rutherford
All published research will turn out to be false.
The problem is ill-posed: can we establish once and for all that something is true? Almost all history had this ambition, yet every day we find that something we believed to be true wasn't. Data isn't encouraging.
My favorite example was a huge paper that was almost entirely mathematics-based. It wasn't until you implemented everything that you would realize it just didn't even make any sense. Then, when you read between the lines, you even saw their acknowledgement of that fact in the conclusion. Clever dude.
Anyway, I have very little faith in academic papers; at least when it comes to computer science. Of all the things out there, it is just code. It isn't hard to write and verify what you purport (usually takes less than a week to write the code), so I have no idea what the peer reviews actually do. As a peer in the industry, I would reject so many papers by this point.
And don't even get me started on when I send the (now professor) questions via email to see if I just implemented it wrong, or whatever, that just never fucking reply.
If the code doesn't work, it seems like a red flag.
It's not an advantage that can be applied to biology or physics, but at least computer science catches a break here.
My favorite for those is to search the code for "todo" and look up that part of the paper. Usually, these are the most complicated parts of the paper. Not always though, sometimes they are just trivial things.
Computer science studies computation as an abstract concept. The work may be motivated by what happens in the industry, but it's not supposed to produce anything immediately applicable. Papers may include fake justifications and fake applications, because populist politicians decided long ago that all publicly funded research must have practical real-world applications. But you should not take them at face value.
Academic CS values abstract results over concrete results, because real-world systems change too rapidly. Real-world results tend to become obsolete too quickly to be relevant in the time scales the academia is supposed to operate.
If you are not in academic CS, you should be careful when reading the papers that you understand the context. Most of the time, you are not in the target audience. Even when there is something relevant in the paper, it's probably not the main result, but an idea related to it. And if you start investigating where that idea came from, it probably builds on many earlier results that seemed obscure and practically irrelevant on their own.
Peer reviewers usually spend a few hours on a single review (though there is a lot of variation between fields). A week would be so expensive that most established academics would have to stop teaching and doing research and become full-time reviewers.
This isn't true. When I'm implementing a paper, I usually go for JUST implementing what they describe, usually by hand. Like if it is a new SQL syntax, I will write a custom recursive descent parser, and hand-roll the query planner, for just the new stuff and hard code some other parts, just as a demonstration. I'm not interested in the industry application part, I'm interested in replicating their work. Once I can replicate it, assuming it is correct, then I will factor the work into a production system.
It's this first part that I am frustrated with, not the full implementation in production software.
Your whole comment reads as some sort of weird gate-keeping where people without the “proper” education could never fully understand a “true” paper. We can, and we do. There’s a reason we learn from truly great computer scientists like I listed above in college, and not the people posting unreproducible work.
My comment was not gatekeeping and more in line with "if you are a mechanical engineer, don't expect to get much value from papers in theoretical physics". "Proper" education or otherwise spending a few years focused on the topic may help, but it's far from foolproof.
The "if you start investigating where that idea came from" part came from my personal experiences. There have been times when I've categorized a paper, on a topic I'm supposed to be an expert, as interesting but practically irrelevant. Only to later find that it became a building block for a major practical result.
If you read a research paper and think you understood it well enough to judge it, you are probably wrong. Even if you are an expert on the topic.
That sounds like utter nonsense to me, but I go through about 5-10 papers per year.
> There have been times when I've categorized a paper, on a topic I'm supposed to be an expert, as interesting but practically irrelevant. Only to later find that it became a building block for a major practical result.
This is usually why I am investigating a paper. I want to understand what they did because I see how it can be used practically, if what they wrote is correct.
I'm working on the wpaxos paper right now. There's even some example code, totally missing some aspects of the paxos algorithm (it also doesn't appear to run in its original form). There's a whole section in the paper about transactions, marked as "//todo" in the code. The flexible quorum implementation is subtly incorrect as well. Anyway, it's a pretty good paper, and I'm fairly certain a large part of it is probably driving cosmosdb (since they work at Microsoft now and I recognize a number of possibilities that echo what cosmosdb can do, along with some same constraints and it would explain the weird billing; but this is a total guess and I have no proof of it). Anyway, your comment is something I totally agree with. Most people lack the required knowledge, creativity, and insight into most papers to truly understand what the value is in the research.
For fun, I am researching magnetic resonance through an iphone's magnetometer to detect stalkers. If it works, I'll have an app (assuming apple doesn't have any issues with it). I wish I could write a paper on it, but I am nowhere near academia these days; which is a whole different shit the academics have done to themselves to make it so only academics can contribute to public knowledge research.
The whole publishing papers thing is broken IMHO, and no, I don't have a solution.
Research papers are published early in the research process. Typically when there are some interesting, likely correct, and possibly relevant results to show, but long before anyone manages to put the results in the proper context.
I speak to a lot of people in various science fields and generally they are some of the heaviest drinkers I know simply because of the system they have been forced into. They want to do good but are railroaded into this nonsense for dear of losing their livelihood.
Like those that are trying to progress our treatment of mental health but have ended up almost exclusively in the biochemicals space because that is where the money is even though that is not the only path. It is a real shame.
Also other heavy drinkers are the ecologists and climatologists, for good reason. They can see the road ahead and it is bleak. They hope they are wrong.
https://pubpeer.com/publications/14B6D332F814462D2673B6E9EF9...
IIRC there was also a paper analyzing how often results in some NLP conference held up when a different random seed or hyperparameters were used. It was quite depressing.
In topics where there is less reliance on relatively small numbers of cases (as is typical for medicine), there is also less reliance on marginal, but statistically "significant", findings.
So areas such as biochemistry, chemistry, even some animal studies, are less susceptible to over-interpretation or massaging of data.
When you get into the humanities a lot of papers "aren't even wrong", as in, the authors don't hold themselves to any standards of logic or rigor to begin with, or aren't even making any kind of identifiable claim about the world, and don't consider rebuttals based on such criticisms to be valid.
A lot of people overlook ecology/climatology because millenarians have made it into such a live wire, but those fields generally have worse problems than medicine. For instance, in medicine going back and retroactively changing patient records for your clinical trial is taboo. It happens remarkably often, but, everyone accepts that it's not supposed to. In climatology they retroactively change climate datasets all the time and if anyone calls them out on it they just attack the critics. They don't even accept the principle that their models should predict data as measured using a constant methodology. It's meaningless in such a field to say "the hypothesis is consistent with the data" because the historical data you analyzed might be replaced with a new version that's fundamentally different.
True vs false seems like a very crude metric, no?
Perhaps this paper’s research claim is also false.
I also saw: a head of design school insisting that they and their spouse were credited on all student and staff movies, the same person insisting that massive amounts of school cash be spent promoting their solo exhibition that no one other than students attended, a chair of research who insisted they were given an authorship role on all published output in the school, labs being instituted and teaching hires brought in to support a senior admin's research interested (despite them not having any published output in this area), research ideas stolen from undergrad students and given to PhD students... I could go on all day.
If anyone is interested in how things got like this, you might start with Margret Thatcher. It was she who was the first to insist that funding of universities be tied to research. Given the state of British research in those days it was a reasonable decision, but it produced a climate where quantity is valued over quality and true 'impact'.
The problem of false scientific claims is global. The UK is by far not the worst offender (mostly places like Iran, India, Pakistan, Saudi Arabia, China are worst affected by paper mills to pick one problem...). And university funding has been tied to research everywhere for decades. It's really got nothing to do with Thatcher.
Really, under an actually libertarian government university research wouldn't be funded at all. From their perspective companies are more than capable of doing research and the patent system encourages publication.
Why do we expect most published results to be true?
I would blame mainstream media in part for this and how they report on research and don't emphasize this nature. Mainstream media also is not interested in reporting on progress but likes catchy headlines/findings.
If we cannot trust that results of research are true, then how can we justify using them to make any kind of decisions in society?
"Believe the science", "Trust the experts" etc sort of falls flat if this stuff is all based on shaky research
Well, don't.
Make your decisions based on replicated results. Stop hyping single studies.
This right here really. The reason people go "oh well science changes every week" is because what happens is the media writes this headline: "<Thing> shown to do <effect> in brand new study!" and then includes a bunch of text which implies it works great...and one or two sentences, out of context, from the lead research behind it saying "yes I think this is a very interesting result".
They omit all the actual important details like sample sizes, demographics, history of the field or where the result sits in terms of the field.
Gell-Mann amnesia is in full effect. "Science" "Journalists" are rarely even one of those things.
The damage has already long since been done. It's great that people are starting to realize the mistake, but it's going to take a lot more work than just saying “stop hyping single studies” in this comments thread to radically alter the status quo.
I once knew a guy who ended his friendship of many years with me over an argument about “safe drug use sites”, or whatever they're called—those places where drug addicts can go to “safely” do drugs with medical staff nearby in case they inadvertently overdose. Dude was of the belief that these initiatives were unequivocally good, and that any common-sense thinking along the lines of, “hey, isn't that only going to encourage further self-destructive behavior in vulnerable members of the populace?” could be countered by pointing to a handful of studies that supposedly showed that these “safe shoot-up sites” had been Proven To Be Unequivocally Good, Actually.
I took a look at one of these published academic research “studies”—said research was conducted by finding local drug dealers and asking them, before and after a “safe shoot-up site” was constructed, how their business was doing. The answer they got was, “more or less the same”—so the paper concluded (by means of a rather remarkable extrapolation, if I do say so myself) that these “safe shoot-up sites” were Provably Objectively Good For Society.
After pointing this out to my friend of many years, he informed me that I had apparently become some flavor of far-right Nazi or whatever, and blocked me on all social media platforms, never speaking to me again.
You're not going to get people like him to see reason by just saying “stop hyping single studies” and calling it a day. Our entire culture revolves around placing a rather unreasonable amount of completely blind faith in the veracity of published academic research findings.
But admitting to the existence of a problem is the first step toward fixing it, and, judging by the downvotes on various comments on this story here, we still have a ways to go before the existence of the problem is commonly-accepted.
The difficulty or risk of using drugs does not appear to be a bottleneck on the amount of it people use. This probably does not hold all over the world, but I'm not aware of anybody actually finding an exception.
As ye sow, so shall ye reap, IRL maybe.
I’m not doubting your claim but I’m wondering how that very weird paper you’re citing bubbles up to the top, when there’s some very middle of the road meta analyses that don’t make outsized claims like access to objective truth.
What do you use to advance your agendas?
I see that you "knew a guy" and apply "common-sense thinking" (that goes against mountains of lived experience, never mind research studies)
I think I have to accuse you of bigotry. Look it up.
That would be a reason to expect those results to be false, not a reason to expect them to be true.
Check out ResearchHub[1], it's a company founded by a tech billionaire that's trying to realign incentives in science
A genius who figured it academic publishing had gone to shit decades ahead of everyone else.
P.S. We built the future of academic publishing, and it's an order of magnitude better than anything else out there.
Do you also say, "Newton a genius? The one who tried to turn lead into gold?"
https://writings.stephenwolfram.com/2024/08/five-most-produc...
Yeah, I think it's fair to judge him by it.
Like all of his work, I thought it was an incredible book, if you just randomly sample 10% of it. I never understood why he doesn't cut more, as he has genius ideas that get really watered down with lots of less relevant details.
I would love if he started doing 1 page tldr's for all of his works.
This is incredible: https://www.complex-systems.com/archives/
"Submissions for Complex Systems journal may be made by webform or email. There are no publication charges. Papers submitted to Complex Systems should present results in a manner accessible to a wide readership."
So well done. Bravo.
There is a lot of academic work that is very obscure and only becomes important later, sometimes decades later, maybe even centuries, to someone else doing equally obscure work, but it always goes somewhere, and the goal is not to "move fast and break things," but create bodies of scholarship that last far beyond any specific capitalist industry or company.
If you published your work online backed by git with hashes with a free public service like GitHub, how could someone steal it?
If you are an academic and don't know git, why can't you pick up "Version Control with Git" from your library or buy a used copy for $5 and spend a couple days to learn it?
> the goal is not to "move fast and break things,"
Who said that was the goal?
Why would you want to remain wrong longer?
If you want to move slower, why not take slower walks in the woods versus adding unnecessary bureaucracy?
It's not about "remaining wrong longer," academics don't care that much about being right. Opinions on works change throughout the years, and its hard to keep track of who did what if nobody can make proper attributions.
>If you published your work online backed by git with hashes with a free public service like GitHub, how could someone steal it?
I'll tell you something, because you have very much outed yourself as a dweeb with this comment: there are physical libraries in the world that are over a thousand years old. GitHub is 16 years old. I would much rather have my work stored in a physical library.
It talks on things like power, reproducibility, etc. Which is fine. There are minority of papers with mathematical errors. What it fails to examine is what is "false". Their results may be valid for what they studied. Future studies may have new and different findings. You may have studies that seem to conflict with each other due to differences in definitions (eg what constitutes a "child", 12yo or 24yo?) or the nuance in perspective apllied to the policies they are investigating (eg aggregate vs adjusted gender wage gap).
It's about how you use them - "Research suggests..." or "We recommend further studies of larger size", etc. It's a tautology that if you misapply them they will be false a majority of the time.
(Edit: spelling.)
I've found the reaction to this article can be pretty intense. We read this in a journal club many years ago and one of the mathematicians who was kind of new to the idea that research papers (in other fields) didn't more or less represent 'truth' said this article was _dangerous_.
Unfortunately the author John Ioannidis turned out to be a Covid conspiracy theorist, which has significantly affected his reputation as an impartial seeker of truth in publication.
> Why Most Published Research Findings on Covid Are False
Well, that's why there was so much focus on replication, multiple data sources and meta-analyses. The focus was there because the assumption is each study is flawed and those tools help extract better signal from the noise of individual studies.
> and that goes against the science politics
I don't think I follow you here. Are you referring to the anti-science populism? That's really the only science politics I'm aware of now that creationism and climate skepticism have been firmly put to rest.
> If only he had avoided the topic of Covid entirely, then he would be well regarded.
I think it's more that his predictions were bad and poorly reasoned and he chose to defend them on right wind media outlets instead of making his case among scientists.
He's not the first well-regarded scientist to go off on a politically-fueled side quest later in his career. Kary Mullis is a famous example.
And please, don't pretend like there is no left wing aligned science politics that is as much based on science as flat earthers. I assume you haven't been hibernating during the covid times. All the doctors who did exactly that are doing fine with regard to their reputations.
I know about Barrington and many of his other claims, but I don't recall him actually saying anything that I would classify as conspiracy theory. Certainly in my world, a credentialled epidemiologist questioning the accuracy of government statistics during a world health crisis, and suggestion that perhaps our strategy could be different, is not conspiracy theory.
I fully agree. (Well, with some caveats. I think credentials matters less than facts. And I think epidemiology is still in its infancy, so I personally don't put much faith in any single epidemiologist.)
Maybe conspiracy theorist is the wrong term. What he did was show a very political concern with public policy (especially IIUC his opposition to lockdowns) and very little concern about the quality of his research or the people it affected.
This article seems pretty decent at containing details: https://www.buzzfeednews.com/article/stephaniemlee/ioannidis...
You mention Barrington, from the Wikipedia article https://en.wikipedia.org/wiki/Great_Barrington_Declaration
> The World Health Organization (WHO) and numerous academic and public-health bodies stated that the strategy would be dangerous and lacked a sound scientific basis.
So I guess maybe less "conspiracy theory" and more "recklessly dangerous" or "abandonment of the Hippocratic oath".
I saw a lot of "epidemiological immune system" activity during COVID- if you didn't toe a specific line, the larger community would attack you, right or wrong. My guess is that this is mainly from historical experience with vaccines and large-scale disease outbreaks, where having a simple, consistent message that did not freak out the population is considered more importantly than being absolutely technically correct.
The Lancet estimated that about 40% of US covid deaths could have been avoided if the administration had better policies. That's a bit over 400,000 deaths. That's about the same number of Americans lost during WWII.
Not all of that can be directly attributed to John Ioannidis's advocacy against lockdowns, but it at least gives us a sense of how big a blunder it was.
https://statmodeling.stat.columbia.edu/2020/04/19/fatal-flaw...
Such careless use of statistics is hardly uncommon; but it's funny to see that he succumbed too, perhaps blinded by the same factors he identifies in this paper.
Beyond that, he sometimes advocated for a less restrictive response on the basis of predictions (of deaths, infections, etc.) that turned out to be incorrect. I don't think that's a conspiracy theory, though. Are the scientists who advocated for school closures now "conspiracy theorists" too, because they failed to predict the learning loss and social harm we now observe in those children? Any pandemic response comes with immense harms, which are near-impossible to predict or even articulate fully, let alone trade off in an unquestionably optimal way.
Ad hominem attacks against ideas can safely be ignored.
(2) I'm not opposed to any ideas in this paper. I think the paper stands on its own merits.
You've tried your best.
>(2) I'm not opposed to any ideas in this paper. I think the paper stands on its own merits.
You're just preemptively setting a limit to how much thinking we can do. After all you made a post in this very thread:
>>So I guess maybe less "conspiracy theory" and more "recklessly dangerous" or "abandonment of the Hippocratic oath".
Which is odd, since medical research has no Hippocratic oath or recklessly dangerous caveats. After all, they were doing gain of function research on coronaviruses in the very city where covid-19 started. Unless geography is now a reckless pseudo science which we must sensor for the good of all.
But John Ioannidis is a physician and has served in a number of medical organizations. Even non-medical researchers are bound by IRB boards for research on humans and more generally are bound by all sorts of codified ethical standards.
It's more accurate to say that his ideas were dangerous and fringe and unsupported by the science. And while he was derelict in the science, he was very active promoting his opposition to lock downs to the White House and to conservative media.
It would be more accurate to say that he heavily fueled the conspiracy theorists rather than he was one himself.
I may have misunderstood your tone, but it sounds like you think it's a good reason to have a bad opinion of him as a person or a scientist, or even prevent his ideas from being heard? I wouldn't want to live in a society like that.
One of the main things that fuelled conspiracy theories the most were draconian measures against dissenting opinions which were perpetrated by social media platforms. Silencing wrong ideas by force damages trust in science much more than engaging with them, and gives these ideas much more credibility.
I don't have a bad opinion of him as a person, although I think he acted dangerously and in a politically motivated way. I think lots of folks were freaking out at the time and their reactions are understandable. I don't expect anyone to be super human. But lots of people were also advocating policies that would (and in some cases did) result in mass death. And I do expect professionals to check themselves and try to prove themselves wrong before they embark on a political mission like he did.
I don't have particularly bad opinion of him as a scientist either. I've known several big name scientist types and usually they're very bright but only really reliable in their established field. You sometimes get to be a big scientist by taking a large contrarian bet, and I would guess he has a natural contrarian streak that served him well in some of his research. The problem with being a contrarian is that you're reactive. Your gradient isn't toward truth it's away from what you perceive as the current central tendency. So you more often end up more wrong than everyone else.
While I don't have a bad opinion of him as a scientist, I do think this episode will make me read his papers much more carefully for conclusions he's reached by contrarian intuition rather than careful reasoning. And it does to me call into question whether his motivation was to find truth or rather to offer a Marx-style criticism of everything to show how much better he is. I don't think the criticize-everything approach has proven productive in the long run.
> One of the main things that fuelled conspiracy theories the most were draconian measures against dissenting opinions which were perpetrated by social media platforms.
I wasn't on social media, so I can't say. It does suck to have your opinion dunked on. On the other hand, social media was full of inorganic influence campaigns. I don't have all the solutions, but I think it's reasonable to have some counterpressure to misinformation.
> Silencing wrong ideas by force damages trust in science much more than engaging with them, and gives these ideas much more credibility.
I'm not sure about this. The campaigns to damage trust in science were quite pervasive and organized. I'd think they wouldn't have spent all that money if the public policies did it just as well without spending on the influence campaigns.
The ideal would be if everyone were educated enough to consume the science directly. But for various reasons mass education is considered political so there's a political divide over who has the foundations to understand it.