Same thing with software security: if software is open source then I still need to trust someone else to determine if the source code is malicious, but it is better than implicitly trusting software because e.g. “a contractor has done a security audit.”
Basically I don’t believe you can truly trust closed-anything, especially scientific analysis. You need _less_ trust when the thing is reproducible, which is the best you can ask for.
The problem is the incentive structure is not aligned with that and so reproduction is never done.
I would instead argue that consensus is the best bar we have to judge science. If the majority of author papers align that means either there is a conspiracy or some fraction of things was reproducibile.
Consensus is a horrible way to judge science. The whole point of basing knowledge on empirical data is so that you don’t need to defer to authorities or consensus (nullius in verba). Consensus based science is just a synonym for Kuhnian paradigms (aka a scientific popularity contest). See Lakatos and Feyerabend for why this is not a logical or useful way to judge science.
Imagine a field where the consensus methods are later found to be in error. In this case the only valuable thing to do is to challenge the consensus so the field adopts better methods. In the meantime, following the consensus guarantees following falsehoods because everyone is wrong in the same way. In many cases the minority heterodox views are the most correct.
Pretending that the tools we have aren't useful because they aren't perfect isn't valuable.
Heliocentric models were good for science even if they were wrong.
You aren't saying anything that moves the line towards a better situation just pointlessly complaining that it isn't good enough.
I agree heliocentric models were useful, but if scientists had followed the consensus methods that produced heliocentrism then we would not have gone past it. It was heterodox theorists who resisted Catholic authority who progressed science. The Vatican are not good peer-reviewers.
If you want a better solution look to Lakatos’ Methodology of Scientific Research Programmes. It is much better than the status quo of relying on peer review. It is rarely used due to inertia of the status quo, but methods like it will become essential as scientific failures mount.
I don’t think my complaints are pointless, we need to take scientific failures very seriously because they eventually become policy failures and cause all kinds of suffering.
Also, I’m not pretending that tools like peer review are useless, I’m arguing that they are useless and providing alternatives that have been developed over decades of epistemological progress.
Basically how can you be sure your difference points to a mistake in the original vs a mistake in repetition.
Additionally I think it was also a little "replication would be cheaper if we did this".
However I think calling it fundamental is a bit of a stress. After all such a paper is going to heavily focus on such details by necessity. You have to categorize to perform such a study and blockers become glaringly obvious.
However IMHO I think that a failed replication caused by inaccurate replication isn't a lost cause and instead acts as a new lens to view the original in. "You forgot to mention whether you controlled for X so I did but it didn't work" is valuable but certainly harder to catch in peer review (where the explicit mention might have caused questions).
Put another way we can replicate as well as we need to if we prioritize it and while better steps are important they aren't as important as actually replicating.
You only once mentioned a single theory which only touches on a part of the replication crisis and not the core issue of "no one can replicate with no funds to do so"...
We put huge amounts of (unpaid) scientific resources towards peer review and it has hardly improved anything. It’s the opposite: “it’s peer reviewed” has become a reflexive defence of published papers with the incorrect implicit assumption that peer review makes the science better.
We should focus on the initiatives that have a proven track record of improving things (like open science) and abandon peer review and reliance on consensus.
As you point out, financial limits prevent replication. So why do we spend so much time and money on fancy journals and reviewers who have such a terrible track record of improving things? We could be replicating instead.
For a recent example of consensus failures, see Nobel Prize winner Katalin Karikó whos seminal paper was rejected by Nature and whos research programme was rejected by her University for being too incremental. Her peers failed because they relied on flawed heuristics like consensus.
What was the alternative (and when?), and what were the results? Who will fund all this replication - basically double-funding all research? Do we want our scientists spending their limited time on old things instead of discovering new ones?
Currently billions are spent on academic publishing even though servers are cheap and the labour is unpaid. This would be a good starting point to fund replications. But that will never happen while the status quo is “peer reviewed = good science”
Also, scientists should be spending time on what they think is valuable. Scientists happily replicate important papers using their existing resources. If no-one cares about your result and won't replicate for free, then pay someone to do it! This is the basis of adversarial collaborations (see FEP collaboration funded by a philanthropic foundation). No-one is advocating that e.g. we replicate newtons law for the Nth time. And how to determine if N is enough? Use Lakatos' MSRP to compare rival theories.
Again, these dreams are predicated on the status quo changing to where scientists care about the content of a paper, not where is it published or that a faceless committee has accepted its foibles. A world where the public response to a non-replicated result is:
"I wonder if any groups plan on replicating?" vs "New paper published in __nature__ has <newsworthy counter-intuitive yet unreplicable result>! Book deal / gov. contract / faculty promotion here I come!"
A lot of obfuscation in general. Poor documentation and omission of critical implementation details means it's way too hard to replicate a paper.
And even if you do replicate there's no way of finding out if the output is wrong because of your implementation or the idea itself is flawed.
Doing proper documentation should be mandatory. I can only talk about CS but yeah more information.
Consensus-building is about crowdsourcing the truth. You reject both authorities and relativism. You believe that, in many cases, the objective truth exists, even if it can be difficult to reach by fallible humans. You are also optimistic enough to believe that experts who share your worldview will accept true arguments and reject false arguments. Eventually, on the average, and with a high probability.
Science _is_ crowdsourcing the truth, by using ideas that other people generate and then testing them empirically to verify them.
This does not require consensus; rather consensus causes problems which make science worse, disagreements are always better for progress. We don’t have to accept current problems when there are better alternatives (see Lakatos for a good example).
Disagreements have a key role in building a consensus. But those who disagree with the consensus are usually wrong, because the world is full of capable people with contrarian tendencies and weird ideas.
The consensus might be correct more of the time, but that shouldn’t be a scientific reason to follow the consensus. You should follow the consensus because you have some empirical reason to agree with it and disagree with it if you have reason to believe a better theory.
The only person who needs to defer to an authority reflexively is someone who is completely ignorant of science. They don’t have the sophistication to judge, but thinking that this extends to sophisticated actors is an error. Consensus is for the ignorant.
This is where we disagree. The ignorant believe in their ability to judge. The deeper your expertise gets, the narrower the scope where you trust your judgment becomes. Because you have seen so many ways things can go subtly wrong. And because you have already been confidently wrong so many times.
1. The scientific method might be reproducible by design, but there can be (and often is) a very large gap between "Science" and what gets published as a paper.
2. Peer review is a grab bag. Sometimes it's obvious the reviewers barely read or understood what was written. Other times they provide feedback that is asinine. Or simply they provide commends and you tell the editor you're ignoring those comments and the paper still gets published.
3. Only a fraction of reviewers (depending on field) will check that the paper includes enough information to be reproduced. This is easy to verify, just go ask different academics how many papers they read that don't include enough information to reproduce. And only a fraction of a fraction make any attempt at reproducing the paper to any degree before signing off on it. (and if they say that they couldn't reproduce it, then see the end of point 2 above)
> If the majority of author papers align that means either there is a conspiracy or some fraction of things was reproducibile.
This is a false dichotomy. It misses some very real dynamics, none of which are conspiratorial.
1. It's hard to publish negative results, especially if the paper written by someone influential in the field. Which means that there's a great disincentive to even try.
2. Competing theories are difficult to study not just because of peer review issues but also because of financial and repetitional issues. Getting funding is significantly harder if your work doesn't use the methods that the funding agencies expect. Same idea with reputation, hard to get your career started if you want to work on theories that go against the prevailing theory or if you have work that contradicts it.
This is the whole idea behind the concept of paradigm shifts.
After all going against the establishment generally results in failure.
Quantum Mechanics and General Relativity may not agree with each other but they both have destroyed mountains of attempts to disprove them.
I am not being rosy I am on the side of "we have to do something" contrasted against "all science is bad".
Not saying you are saying that but some posters are.
We can certainly improve things on many vectors but throwing away all existing science isn't the way forward either.
I agree. Unfortunately for all of us the rest of published literature is nowhere near as battle test as these two theories, nor would they survive such testing.
> Not saying you are saying that but some posters are.
I didn’t mean to imply that you were. But the it doesn’t sound like the personally you replied to was on the side of “all science is bad” either.
I am closer to the “much ‘science’ is junk”. There is much good work a being done, but the garbage is very much there, more so in some fields than others.
The null hypothesis at this point for most new papers one reads is certain fields (looking at you nutrition & health) is “this won’t replicate”. It’s certainly also true for much work primarily involving modelling; the null hypothesis as a read should be “if this doesn’t have code, I won’t be able to reproduce/replicate it”.
We should not throw the baby out with the bath water, but we should be frank about the current state of things.
I have suggested we throw away a part that is causing problems. You seem to be rushing to defend scientific failures by saying “it doesn’t matter that a lot of studies don’t replicate, we should follow the status quo regardless.” We can throw out the status quo and keep (the good parts of) science intact.
But, as you also say, it’d be a Really Good Thing if it was about reproducibility! Imagine a world where instead of some people writing an essay, their “peers” giving it meaningless comments, and some editors at a paper selecting it to be enshrined as “valid”, we totally flipped the script:
People perform a scientific observation. They record their methods and results, and put it into some freely accessible store of data regarding the question at hand. Anyone is free to consult the store for any question, and observe how many entries it has, and how their results compare. If an entry has very few results, the person consulting it with the question would be encouraged to create a reproduction of their own, and share the results they derived as a sibiling of the original paper.
I couldn’t understand the last couple paragraphs. sarcasm?
As for the last few, I’m basically just saying we should dismantle the current “journal” concept entirely (it stands only to benefit those who receive the fees it takes and those who derive self-worth from being published in a “prestigious” journal), and replace it with a system by which for any given scientifically testable hypothesis, a collection of many different reproduction attempts and their respective methodologies and results are immediately available all side-by-side. With that in place, no scientific result would derive any credibility from being “peer reviewed” or not, but rather from the quantity and diversity of reproduction attempts it has faced.
This database should be free to query and free to insert into. Individual papers may support community comments to serve as the weak “peer review” we currently have, but at no point should these comments be considered anywhere near equivalent to a full reproduction attempt.
In my field replication usually involves a lot of guesswork since papers don't usually have enough information to be certain of their exact methods (no code or data unless the planets align or the authors are friendly) and yet peer reviewers happily apply their rubber stamp.