Refuted papers continue to be cited more than their failed replications
statmodeling.stat.columbia.edu
statmodeling.stat.columbia.edu
An alternative to GoogleScholar is Research Rabbit, an AI enabled platform that maps the networks between paper allowing for an author to see who cited the paper they intend to cite, who did the original paper cite etc etc.
The root of the problem is that it’s impossible to build a scientific career out of publishing negative results (assuming you can even publish them at all). What actually matters is that your original publication gets into a high tier journal and gets good press. Once that’s done, no one cares if what you published was garbage or not, because by the time anyone figures that out (if they ever do), you’ll have tenure, big fat multi-year grants, and will be out on the speaking circuit.
I think the tradeoff we have taken is "lots of science, mostly good" vs "little science, 100% correct". Similar tradeoff is taken with software & wikipedia, for example.
Or do you apply the same rigorous standards to, say, software? Should every piece of software that is publicly distributed be thoroughly reviewed for correctness?
The fundamental issue: what is in it for the reviewers in this case? Scientists are people, and you can't really legislate interest in replicating other people's work that is unrelated to your own, unless you really spend money on it.
And there's some bad news on that front: https://www.science.org/content/article/final-u-s-spending-b...
It's flawed for sure, but there's so much around you that comes from this system if you care to look for it.
For example we can make no certain conclusions on the "Hard question of consciousness." As a result there is nothing concrete about the concept of consciousness because it's not even a well defined enough question. Anyone attempting to make a claim are making it in an epistemic vacuum. So it doesn't matter who is talking about it, the state of knowledge of the concept of "consciousness" is so lacking that it doesn't even make sense to discuss.
In my view most of the persistent problems in humanity are in this class: we haven't defined measurements for the issue in question well enough to actually be able to approach a solution.
I always keep in mind that: "I'm probably wrong about my position"
If you are seeking data to confirm your idea then you'll be wrong more often than not
If you are seeking data to falsify your idea then you'll be less wrong more often
If we stop replicating those limitations, it's not unreasonable to expect that the social problem will change also.
I don't think the answer is search engines in the traditional sense, but we probably do need something that knows how to search. If whatever we use to view research were to display warnings wherever a retracted citation was viewed (or whenever one the citations' citations were retracted...), similarly to how we display SSL certificate expiry warnings in a browser, I think that would create a dampening effect on the momentum that a retracted paper can have.
"Retracted" probably isn't the only color we'd want here. "Replicated" might be another. My point is that summoning up-to-date metadata on a something published last year should just be the default view mode, not something you have to ask a grad student to go do.
Knowledge graphs, not dead trees.
This is called technological determinism and is not an accurate heuristic for how humanity adopts tools
Humans generally adopt tools broadly after the social environment allows it, not when the tool is capable. This has been shown over and over. Never has there existed a tool whose introduction was immediately and universally adopted.
There’s almost always a period of introduction, then decades of middling adoption and refinement, then the social climate changes and production adjusts enough to adopt en masse.
This was true for every major invention and is an artifact of human social structures
It's heavy handed, but I think every time a study fails to reproduce its results, every single published paper that uses that refuted study is notified, the authors get 6 months with a warning banner on top of the article, and if a correction is not submitted, it's automatically removed.
Science isn't supposed to be easy. We need a system that prioritizes high quality output, not volume.
so rather than compare total number of citations, you want to see that:
- citations per year drop once refuted
- if citations continue after refutation, they co-occur with the refutation study
To analogize to the legal field, which they mention in the article, if a case I cite is good law on one point of law, but no longer good law on another I may cite it as _Plaintiff v. Defendant_, 123 F.3d 456 (1998) (reversed on other grounds, _Defendant v. Plaintiff_ 78 S.Ct. 910 (2000))
Just because one finding in a paper wasn't replicated doesn't mean that the paper can't or shouldn't be cited for other findings.
So even for the small percentage of people who have some narrow focus of inquiry (usually at work) they only have the desire or ability to evaluate the first or second order inputs and effects.
In general this isn’t a problem for day to day life. The packaged chicken on the shelf fits perfectly into recipe, and there is little to no thought of any other externalities (what systems am I “voting with my dollars” to support? What is the relative difference in supply chain between two options)
However if you are going to try and align your conceptions of how systems/environments work and what you desire for future state, then you do have to do the system decomposition and that’s actually non trivial.
So there’s no real “market” for epistemological rigor outside of a small group of people who come across as unreasonable in their demand for rigor, so this is the likely result of a system that is built for personal or organizational success and not an ego-free search for “truth” via intentional rigor.
Luckily the small group continues to grow as a total number
We seem to have a lot of hypertext systems that keep track of citations. Seems like it would be cheap at this point to put a black mark on papers that don't replicate, and to propagate that black mark to every paper that references the bad paper without also referencing the failed replication. Put a big red X over the abstracts.
Does it? The rate of paper production is so high that it has become a logistical problem in and of itself and wild goose chases are common. Surely some of that energy could be profitably redirected into replication and consolidation.
The black mark system could work but seems like a recipe for some nasty politics. I tend to think that baking replication into the culture might be a better approach -- science advanced beyond secret-hoarding when information sharing became institutionally valued, and I believe that it could advance beyond p-hacking if replication became institutionally valued. Fund & cite replication efforts and the rest will follow.
Yes and no. The scale of the checking problem is formidable, as you're aware. However, many claims in scientific papers are weak. Even highly cited papers can contain nonsense that doesn't stand up to mild scrutiny. Given that, I think progress can be made towards the problem if every researcher simply spent more time checking previous work. Checking everything is impossible, but progress can be made by prioritizing, dividing the work up among everyone, and doing checks that don't take much time.
As an example of a highly cited paper with easily noticed major issues, consider this comment I posted on PubPeer: https://pubpeer.com/publications/95455FA4147A9CBD5EAA5185D21...
The first version of this paper was published in 1991 and has over 300 citations. It's still cited to this day. Despite this, in 2016 I read the later journal version of the paper and pretty quickly realized that the first part of the paper was basically nonsense. You don't need to look hard to realize that either! The problem can be found by simple spot checks, looking at the trends of the equation. (I can't quite recall, but I think that's how I figured out that the model was wrong in the first place.)
My paper refuting the model to date has zero citations that aren't from me. The closest I've received so far was an email from someone who was curious that I was apparently the first to notice the problem, over 20 years after the first publication of the model.
Rigorous knowledge in any domain is both difficult to obtain, and it usually commands a small premium compared to the other things you could have sunk your time and attention into. I see at least two options for getting around this:
- Wirehead yourself into thinking truth is beauty and beauty truth, and then tilt at what many others will consider windmills (the fools!). At its best this is probably what folks like Isaac Newton were up to.
- Go all in on the profit motive, embrace competition as a refining fire to force you to be more rigorous or go broke trying. This reduces the problem to knowledge being "just" really hard to obtain.
There may well be more, those are just the two most obvious ones to my eyes.
[1]: https://bazhum.muzhp.pl/media/files/Studia_Humana/Studia_Hum...
It’s pessimistic at the extreme and throws one’s hands up to nihilism
I prefer pushing through that nihilism and into realizing that absurdity of the universe is not a fixed barrier but a challenge to solve for the entirety of the universe
I'm not, so I find the message of finding niches people haven't pushed hard enough into to make the world a better place (and maybe make some money while you're at it) to be pretty uplifting.
Bloodletting is well understood that it doesn't work but there was a time where the most trusted people performed that
Why is that even possible? There was absolutely no epistemic proof that bloodletting was associated with a falsifiable hypothesis which would lead to improved health outcomes - in fact it was all based on these "humors" in the body that needed to be "balanced" which makes no mathematical sense whatsoever.
Even then, as early as the 1400s there were scientists saying "This is actually bad" - and yet George Washington had bloodletting done on him in his dying repose at the end of the 1700s
300 years of an activity that is bad for you that was encoded into medical practice
The conclusion we should come to is that MOST of our assumptions about the structure of the world and how to optimize within it, are based on literally nothing but aggregated social velocity, instead of claims that are measurable and testable
1: Refusing to make a firm affirmative statement on an issue, when that's what people want
2: Correcting what you have claimed in the past, which makes you look gullible and indecisive
With the example of bloodletting, people wanted to believe a doctor knew how to cure them, and once somebody had done bloodletting even if they had doubts later, they would look like fools if they convinced people they had been harming their patients all along.
It's easier and often more rewarding to be a fraud.
we have created systematic incentives for fraudulent behavior
If we get rid of those incentives, then that will expose all of the frauds, and so as a result, they have no incentive to overturn the system