Propagation of mistakes in papers
databasearchitects.blogspot.com
databasearchitects.blogspot.com
> Millikan measured the charge on an electron by an experiment with falling oil drops, and got an answer which we now know not to be quite right. It's a little bit off because he had the incorrect value for the viscosity of air. It's interesting to look at the history of measurements of the charge of an electron, after Millikan. If you plot them as a function of time, you find that one is a little bit bigger than Millikan's, and the next one's a little bit bigger than that, and the next one's a little bit bigger than that, until finally they settle down to a number which is higher.
I once did a web-of-science search for citations to a foundational paper in my field. It was published in volume 13 of a particular journal, and that was listed in a little over 90% of the citations, but the other citations all listed the journal as 113. My assumption is that somebody cited it in error, and that others were basically copying the citation from the bibliography, rather than going back to the original paper to get the original metadata.
Does this mean that about 10% of writers were basically lying about having read the original paper? Well, maybe. But I fear that the number might be higher than 10%, because the correct citations might also have resulted from just copying from a bibliography.
I tell this story to my students, in hopes that they will actually read the original papers. Quite a few take my advice to heart. Alas, not all do.
Primitive science (or even pre-publishing science) doesn't get cited because humanity figured it out before our current system was in place.
It may sound silly, but no one feels the need to cite Eratosthenes when implying the world is round.
But many people do feel the need to cite the colorimetric determination for phosphorus (an SCI top 100 paper) even though it was published 100 years ago and is generally considered “base-level science.”
It is certainly an interesting paper to read, but I’m not sure I need every scientist to read it in order to believe they know how to do colorimetric analysis.
Only for some categories I expect the authors to have read the cited paper. Some of them don't necessarily mean the cited paper is high quality. Some are recommended reading to understand the original paper, some are'nt
https://www.gwern.net/Leprechauns#citogenesis-how-often-do-r...
Your 10% isn't far off from the 10-30% estimates people get, so not bad.
Namely, paper references always reach back in time. Papers don't reference papers that were written after they were written. And if that sounds stupid, bear with me a second.
We've talked a lot about the reproducibility problem, and that's part of propagation errors in papers (I didn't prove this value, I just cribbed it from [5]). If we had a habit of peer reviewing papers and then adding the peer review retroactively to the original paper, both for positive and negative results, would we slow this merry-go-round down a little bit and reduce the head-rush? Would that help prevent people from citing papers that have been debunked?
You're modifying the thing so that future ∆p^{i+k} are added to the delta-mapper so that ∆p is appropriately modified accounting for that ∆p^{i+k}. It's like path-compression in a union-find structure.
It is interesting as a helpful approach but does suffer from the pingback spam problem, right? And I have a slightly sneaking suspicion that it is not an accidental oversight in science that leads to these problems.
Actually, the opposite happens quite regularly: the author X got a pre-print of a paper by author Y. From Y, he knows that the paper will be published in the future in an upcoming volume of a scientific magazine. So, X already writes this future reference into his paper.
I've speculated before that peer review gives researchers false confidence in published results [0]. A lot of academics seems to believe that peer review is much better at finding errors than it actually is. (Here's one example of a conversation I had on HN that unfortunately was not productive: [1].) To be clear, I think getting through peer review is evidence that a paper is good, albeit weak evidence. I would give the fact that a paper is peer reviewed little weight compared against my own evaluation of the paper.
I think this depends on how you define good. I'm sure there's some variation across fields, but peer review generally seeks to establish that what is presented in the paper is plausible, logically consistent, well-presented, meaningful, and novel. That list is non-exhaustive, but correct is very hard to establish in a peer review process. In my experience, it would be rare for a reviewer to repeat calculations in a paper unless something seems fairly obviously off.
As a computer scientist, it would be even more rare for a peer reviewer to examine the code written for a paper (if it is available) to check for bugs. Point being, there are a lot of reasons a paper that appears good may be completely incorrect. Although this is typically for reasons that I as a casual reader would be even less likely to distinguish than a reviewer who is particularly knowledgeable about that particular field.
Peer reviewers are not monitoring how experiments were conducted, they only have access to a data set that is by necessity already highly selected from all the work that went into producing the final manuscript. The authors thus bear ultimate responsibility.
When considering published work close to mine, I use my own judgement of the work, regardless of peer review or which journal it is published in (for example it may be in a PhD thesis). For work where I am not so familiar with the methodologies, I prefer to wait for independent verification/replication (direct or indirect) from a different research group, which ideally used different methods.
And the Scheuermann and Mauve paper mentions that they picked the value (0.775351) from the Philippe Flajolet paper that only mentions it without the extra 5. It's not that it was calculated again, reviewed or something like that. It was simple picked up and typed wrong.
The assertions I've imagined would be among the innovative papers on the established papers values.
One outcome is that this would strongly encourage fitting in and discourage innovation and disruption.
It could discourage publishers from correcting a past error even stronger?
So I'm not sure if I like the idea at all to be honest.
Let's keep it uninvented as it is.