[1]: https://www.experimental-history.com/p/science-is-a-strong-l...
[1]: https://www.experimental-history.com/p/science-is-a-strong-l...
I don't think that this is true at all. Weeding through bad papers is, at a minimum, an opportunity cost, as is a good paper built on top of a bad one.
Also, there is a societal cost in that bad research can get picked up and believed by people, like the anti-vax crowd. Or, bad research can be used to push an agenda, like anti-climate change.
You shouldn't have confidence in a paper because it was published in Nature, you should have confidence in it because it was published by Harvard or the University of California or Google Research, who puts the name of their institution on it and thereby stakes their reputation. For the preeminent researchers in a field, their own names may be enough for people in that field to trust the result.
You can still have a paper published by nobody, and if the results are interesting, researchers at known institutions can try to replicate it. Until then it's nothing. But if they do, the nobody gets cited and credited with being the first and takes a step to building their own reputation, and the known institution gets credited with knowing when to spend resources following up on something significant.
You can have papers published by Google Research or by nobody now: turns out academics like the journals' attempts to curate interesting stuff more than only ever reading stuff on their bookmark list or that's mailed to them in a desperate attempt to get some engagement. And for all that they might not like the long delays and low chance of success associated with submitting to Nature, they like the idea of trying to persuade an academic at an elite institution to go to the trouble of replicating their research as the only means to build their reputation even less...
Their business model isn't to put the research behind a paywall, which is the relevant distinction.
> turns out academics like the journals' attempts to curate interesting stuff more than only ever reading stuff on their bookmark list or that's mailed to them in a desperate attempt to get some engagement.
This is what conferences are for, and then you have the conference organizers curating the talks but everyone gets the list of talks and links to all the freely available papers even if they can't afford to attend the conference.
how many such reputations can you keep track of and gauge trust along?
and how is reputation gauged if peers don't review work?
Why not? The same is true of an individual researcher. But if they do, it damages their reputation, so they have an incentive to not. The same as the institution.
because it can produce bad researchers, and thus cannot be trusted to indicate good vs. bad researchers
thus, it is a bad proxy for trust of individual researchers
>The same is true of an individual researcher
individual researchers are not research institutions, so the same literally cannot be true: one employs the other, the inverse is not true
> if they do, it damages their reputation, so they have an incentive to not
empirically speaking, they objectively do, and their reputation is not damaged, and thus above proposition about reputation and incentives does not appear to be true enough to stop it from happening
recall the topic is how to gauge individual researcher reputation in the first place. Either we do it on an individual basis or a group/heuristic basis, and of the latter, research publications are a better proxy than what school one went to, but the former is better than both
That's actually a great point. Who says this isn't already the case? There isn't much evidence to suggest that things have gotten better after the Why Most Published Research Findings Are False paper in 2005. I think people should be extremely skeptical of anything they read, irregardless of a peer review stamp.
This is a mistake that many HN readers make; they think that if one is equipped with some above-average level of intelligence, one can discern the validity of new research. But this is wrong. It usually takes many years of study before one can begin to clearly understand what is even being said, let alone whether it has any veracity.
People of above-average intelligence often resent this fact, because it suggests that their smart opinion isn't as valuable as the opinion of an expert, but that is simply a sad fact of life. If they haven't put in the work in that field, then they don't know what they're talking about in that field, and they are incapable of applying skepticism in that field correctly. This is precisely and unfortunately why we are forced to place our trust in experts.
As an example of an issue that doesn't require much expertise to spot, I was reading a health paper a couple of days ago that reported on an intervention in a specific population. The paper said words to the effect of "[the intervention] was effective, especially for men". The associated chart not only had (somehow?!) got a legend that didn't match the actual chart lines, but whilst the line for men did indeed go down the line for women was very clearly flat! This should have been reported as "no effect in women but an effect in men" but wasn't. When the wording doesn't match the data being reported, that's a good sign in any field that the researchers know they're treading on thin ice. That particular claim was a correlation/causation fallacy anyway, which is very common in health. They didn't have any proof it was their intervention causing the reduction and there were a bunch of reasons to suspect it wasn't. But the intervention was long term and high effort so it's not a surprise they wanted to find something.
To gauge what sort of p-value suggests a useful model requires familiarity with the area and the broader context of the research.
If a field accepts P=0.05 as significant it means 1 in 20 results can be false positives just by random chance. Now think about how many papers get published, and many of them report more than one thing. The right threshold should really be an order of magnitude lower. It's not OK for scientists to report FPs at that rate, and that's a big part of the reason for declining confidence in science.
That is not the correct interpretation of a p-value. See the ASA's statement on p-values.[0]
>Researchers often wish to turn a p-value into a state- ment about the truth of a null hypothesis, or about the probability that random chance produced the observed data. The p-value is neither. It is a statement about data in relation to a specified hypothetical explanation, and is not a statement about the explanation itself.
Moreover, while it's fine to suspect that p-hacking took place with p-values just under 5 percent, it is merely a suspicion and nothing more. Throwing out p-values you don't like without further evidence is a different sort of violation of the methodology; it is like p-hacking in the other direction.
>The right threshold should really be an order of magnitude lower.
This is false. Again, see the ASA's statement.[0]
>Practices that reduce data analysis or scientific infer- ence to mechanical “bright-line” rules (such as “p < 0.05”) for justifying scientific claims or conclusions can lead to erroneous beliefs and poor decision making. A conclusion does not immediately become “true” on one side of the divide and “false” on the other. Researchers should bring many contextual factors into play to derive scientific inferences, including the design of a study, the quality of the measurements, the external evidence for the phenomenon under study, and the validity of assumptions that underlie the data analysis. Pragmatic considerations often require binary, “yes-no” decisions, but this does not mean that p-values alone can ensure that a decision is correct or incorrect. The widespread use of “statistical significance” (generally interpreted as “p & 0.05”) as a license for making a claim of a scientific finding (or implied truth) leads to considerable distor- tion of the scientific process.
In short, these things often require a strong statistical literacy to interpret, which most people complaining about p-hacking do not possess.
[0]https://amstat.tandfonline.com/doi/pdf/10.1080/00031305.2016...
Yes, obviously the whole notion of a hard threshold is a bit nonsensical to begin with, but as the ASA statement says "Pragmatic considerations often require binary, yes-no decisions". There doesn't seem any way around that. People face decisions like, shall we continue to fund this investigation? Yes/no. Should we recommend lifestyle changes to the public? Yes/no. It doesn't make sense to try and map a P value into a budget, for example. At some point you need a threshold (and likewise for effect size and other things). Of course at some level there is fuzzyness, which is why I said if a paper is mostly reporting 0.049 values then ... and I didn't specify what, exactly because the conclusion should be something like "fuzzily suspicious and should look closer".
So it's fine for the ASA to complain about statistical significance leading to "distortion of the scientific process", but their proposed alternative can be boiled down to doing more peer review, as most of the things they tell researchers to look at are things researchers will never conclude in the negative about their own results, like study design (because if they did they wouldn't have got to the point of calculating a P value to begin with).
Finally, I disagree when you say "That is false". Scientists aren't going to stop using statistical significance thresholds because they do need to make binary decisions at some point, and dropping the threshold they currently use by 10x would immediately yield major improvements in replicability and robustness.
Anyway, put P-hacking to one side if you dislike that discussion, it's fine. It's not actually the thing that bothers me the most when I read papers. A big gap between the prose summaries and what the data [analysis] actually shows is a much more common source of distortion, IMO.
This is precisely the sort of misunderstanding that surrounds p-values, and it is what the statement aims to correct. Lowering the standard p-value threshold does not imply an improvement in these factors. The problems of p-hacking and creating false positives lie in the disclosure of the methodology, not the threshold of p-value. That is the whole point.
This will be my last comment in this chain. I'm just trying to clear up some persistent misconceptions that I see on the internet.
In many cases results can easily be independently verified. This is why it works for AI. If you publish a result, it should come with code anybody can run. If the code doesn't exist or doesn't do what you say it does, you're a fool and everyone can ignore you. If it does, you don't need anyone's stamp of approval to prove it.
But that doesn't work with medical trials or things of that nature where independently verifying the claims is expensive.
Responsible media would act as a gatekeeper. The problem is, most media utterly gutted scientific journalists for more profit, so they completely lack the basis to evaluate and supply context on research. Others, particularly boulevard media, willfully ignore any kind of ethics for clicks.
On top of that comes a general media illiteracy and media distrust that makes it even harder to combat because there's an awful lot of media that thrives on intentionally pushing crap to people.
How does one come across these bad papers in the first place, such that they must be weeded through? Maybe we need a different method of curation such that they stay below the noise floor.
There's a popular argument about freedom of speech that comes down to "the solution to bad speech is more good speech, not limiting bad speech," but that only works (assuming it's true at all) when there's some way for a listener to discern truth with some research or effort.
> "Good papers are good because they have good arguments"
No they are not. Good papers are good when they collected data correctly and then presented the results fairly. Arguments about what that data means hardly enters into it. You can't argue whether the data is correct or not without gathering it yourself, and you can't gather it yourself.
Yes exactly. If it is too hard for a group of scientists to figure out a flaw, then it is also too hard for two peer reviewers to figure out a flaw. Only time can detect all errors in science, assuming that nonconformist papers are allowed into the conversation.
> Good papers are good when they collected data correctly and then presented the results fairly.
This is only true for empirical papers. Good data is a part of a good argument in the case of empirical papers.
The issue isn't that flawed papers are necessarily hard to spot, the issue is that indiscriminate publishing means that anyone studying a topic has to wade through a lot of flawed ones before they find anything halfway useful. Time and attention are finite. Peer review and journals caring about reputation, in theory, caps the number of readers of rubbish papers at 2 (and disincentivises writing unpublishable crap to submit in the first place... although this is now eroded since generating rubbish papers is now effort-free)
You actually can argue about the data being "correct".
The way to do is, is studies need to be much more radically open with their raw data sources.
In AI, that would mean actually releasing the code you used, so people can try it themselves.
Or, in more sociology stuff, where you are doing, I don't know, interviews with humans, you could release the physical videos of your interviews.
And then for science stuff, show pictures/vidoes of the science you are doing.
I'm sure there would still be way to hack this. But significant more code and data transparency would do a huge amount for allowing other people to replicate or verify your work.
Instead of targeting peer-review, i recommend approaching this problem as one of incentives. Under the current system, what is the incentive for a reviewer to read and critique a paper? None. If they were paid in cash, there is a clear incentive. The publishers do not want this overheard and the associated legal requirements so they seek out "volunteers" and "compensate" them with rubbish like book copies and "reputation". The fault is with publishers, not the peer-review process.
At the moment, the current model allows parasites like Elsevier to derive the maximum benefit while not paying most folks involved in creating the actual value - the reviewers, editors and authors.
Citizen journalism tried this, failed and is now hollowing out established journalism. We need proper funding and proper incentives for journalism just as we need proper incentives for academic research.
Many universities limit the number of hours their faculty can commit to paid external activities. Paid peer review would be one of those activities, and it would have to compete against other activities. Such as consulting, which can be very lucrative in some fields.
So maybe you have to pay $10k peer review fee when you submit a paper. And then it gets rejected after the first round of reviews, because you aimed too high or the editor and the reviewers just didn't like the topic. You resubmit to another journal and pay another $10k. After a couple of additional rounds of reviews ($5k each), the journal seems to be interested in publishing the paper. But reviewer 2 wants you to cite some of their papers that are not really that relevant. Do you agree, or do you argue against and risk another round of reviews (another $5k)? Or maybe you get the feeling that reviewer 3 is stalling the process with superficial requirements, as they get easy money from the reviews.
Except that most academics can't afford that. American academics are massively overpaid by global standards, and peer review would naturally be outsourced to developing countries, where the expectations of pay are more reasonable. Many of the academics there are reputable, after all. Unfortunately the institutional culture is often more problematic, and corruption also tends to be more prevalent. Do we really want to let those institutions shape the practices of science worldwide?
One of the unfortunate features of capitalism is that being a middleman is more profitable than doing the actual work. Instead of a system where you pay $x for reviews, we could end up with one where you pay $1.1x to a middleman, who then pays $0.8x for the reviews. The middlemen get even richer than in the current system, because there is more money in publishing than there used to be.
At the moment, greed drives open access publication fees. Journals charge thousands of dollars because they can. There is no logical justification of the cost. There are several journals that charge hundreds of euros for review although they do not pay reviewers.
Note also that not all review needs to be done by academics. folks in industry can contribute equally but their employer pays for their time so that boils down to altruism or business priorities. A reasonable payment for time creates an incentive for the world beyond academia and this is a net positive. The payment does not need to be prohibitive, just enough that the reviewer has an incentive to do a good job and isn’t pressured to rush. The problem thus is not limited to the journals or the peer review process but inventives.
A small payment like $200 can be more of a disincentive than an incentive. You have to deal with bureaucracy to get paid and more bureaucracy to pay taxes for it. And if the money comes from a foreign source, amount of the compliance bureaucracy can be absurd. Getting paid is simply not worth it, if the payment is too small.
If by progress you mean a constant creep of SOTA on meaningless benchmarks, then yeah, we have "progress", but that progress is a strange kind of progress that is measured only on its own, self-selected criteria, and that does nothing to advance the total sum of knowledge. The main beneficiaries of this "progress" are large technology corporations who have now taken over research in AI and are turning it to profit. There has definitely been proress in money-making schemes and personal aggrandisement schemes of charlatans and mountebanks, in AI, aplenty.
If by progress you meant scientific progress, then, no, there has been none of that in recent years. It is questionable if there has ever been any kind of scientific progress associated with AI. While early pioneers of AI, like John McCarthy and Claude Shannon, wished to establish a scientific programme of research in human and machine intelligence, with the purpose of understanding the former by developing the latter, this programme was soon set aside, and activity concentrated instead in what McCarthy used to call the "look ma, no hands disease of AI":
>> Much work in AI has the ``look ma, no hands'' disease. Someone programs a computer to do something no computer has done before and writes a paper pointing out that the computer did it. The paper is not directed to the identification and study of intellectual mechanisms and often contains no coherent account of how the program works at all.
http://www-formal.stanford.edu/jmc/reviews/lighthill/lighthi...
And this is where we still are today. AI is no science.
FWIW I'm also really critical of peer review/journals, just don't think this is a great line of analogy.
In that sense, peer review should be an equalizer, because the only way to be read and cited then is to have the holy academic lineage, which is what would happen if we abolished peer-review as it exists in the current academic word we're in.
No one has the time to keep up with all the research coming out, not to mention all the fake research published in bad faith. Life is not restricted to Elsevier and ArXiv, and there are sane alternatives out there.