Nontrivially fillable gaps in published proofs of major theorems
mathoverflow.net
mathoverflow.net
My view is that peer review is a fairly low (and random) bar to pass that doesn't necessarily say much about the validity of a work. More people should check things more carefully.
Recently I got reviews for a paper I wrote back, and a reviewer was skeptical of one claim I made because (paraphrasing) "If this were true, that would mean all previous researchers were wrong." They're exaggerating as it would only mean many researchers were wrong; I think some were skeptical of this for a long but kept quiet about it. But the idea that "everyone seemed to accept it, therefore you are probably wrong" independent of any arguments made seems popular, despite it being obviously fallacious. I have some arguments. The reviewer should engage with those arguments rather than make some appeal to popularity. Edit: To be clear, what I was claiming here would overturn something dating back to the 1930s that last received a well-accepted revision about 40 years ago.
I think fixing the two are almost one and the same. Peer review without replication is just a place for ivory tower intellectuals to review from an armchair.
Replication helps to demonstrate that other people have not only "reviewed" the work but have gone through the trouble of setting up the problem and re-identifying the (potentially undocumented) pain points.
I do think, however, it may not be tenable in the current structure. I think there is too much cost (both monetarily and opportunity) to expect replication to be part of the peer review process. But somehow, we still need to incentive replication.
If I'm a reviewer, then that means that I'm asked to volunteer a bit of my spare time to facilitate the evaluation of my peers by helping the editor (who's expected to be competent but not knowledgeable in every niche) to read the paper, and answer the following questions: (a) is this topic relevant and novel in my sub-niche; (b) is the content adequately described in a way that the target audience (i.e. people in my sub-niche like me) can understand clearly; and (c) does it have any obvious flaws or ignorance of things and previous research that are (or should be) well-known in my sub-niche. So I provide my niche-expert opinion about these questions to help the editor make their decision on which papers to include and what changes to mandate.
That is peer review - in some sense, establishing whether the author is "behaving as a peer". Replication is something that comes after the study is published. If the system would require me to replicate your experiment for a review (assuming that I'm interested enough in the particular topic - I review papers that are tangentially interesting, but not the direction that I'd want to do myself), then that would be unreasonable to have it done as a volunteer service, it would require months of full time work to make a review, and who would pay for that work? If the authors would have to fund a replication before publishing, that places an enormous barrier on publishing and, frankly, I would want to read results that have not passed that barrier because I can't assume that everything that's useful for my research will be able to pass it. As the reader in the target audience for scientific publishing (i.e. other researchers actively working in that field - scientific journals and conferences are made by a niche of scientists for that niche of scientists, it's how we communicate between ourselves, disseminating knowledge to outsiders is not the primary goal for the vast majority of academic publication venues, that's what textbooks and monographies are for) I want authors to publish their results without a prolonged vetting, because I'd prefer to read them sooner rather than later.
How would moving peer-review to after publication enhance trust? It seems to me one of the problems currently is the vast amount of publication. Much of it seems derivative or of little value other than resume padding. While I agree the current peer review system isn't ideal, it does seem to at least provide some throttling of the volume of publications.
Academic publishing is roughly where the music industry was shortly after Napster. There is no reason they should get away with charging crazy fees for putting a PDF up on the web and people are sick of it. Journals that don't accept the new reality of free information wont last long. Once the papers are freely published, peer review will follow suit and happen out in the open like comments on a blog today.
[1] https://academia.stackexchange.com/questions/11097/ethics-of...
The linked question is interesting for me to think about given one of my relatively recent experiences (this is all in the area of statistics, so kind of directly relevant). I was asked to review an article, and as a counterpoint to one of my concerns, the author cited an in-press paper in a fairly well-respected journal. So I go to look at it, and it makes little sense to me, in that it contradicts a bunch of other things that are known. I look at it closer, and it turns out there's a subtle but important notational error carried through much of the proofs that basically invalidates the whole paper.
So I contact the editor of that journal to feel out the response to writing a commentary about the issue, with a fairly detailed explanation of the flaw in the proofs. Instead of being receptive or at least neutral, the editor throws up all sorts of obstacles — not exactly threats, but strong discouragement in the form of a long list of criteria that had to be met, several of which were completely unnecessary. My colleagues (who are also editors of other journals) got even more upset than me, believing that the editor was trying to bury the error and so forth. One even threatened to expose the exchange on twitter or something.
The truth is, I don't know that I cared that much about this particular topic to really put the effort into writing a commentary, fighting with the editor to publish it, and skewering the author's ego in the process. My friends and I discussed just putting the commentary on an online archive, but it wasn't clear anyone would make the connection with the paper. Also, for unrelated reasons, I shortly afterward wrote a different paper that was more comprehensive in scope that sort of superceded that flawed proof anyway (that is, if someone read this paper of mine, the results of the first erroneous paper would probably seen as as irrelevant).
In this case, the flawed proof/paper was peer reviewed. Making the reviews transparent wouldn't have mattered because in the end the paper was published. Maybe one of the reviewers raised the issues but it was published anyway, so published reviews would tip a reader off? But at that point where are we? If I had done something (and maybe I still will?), I would have just published it in a public academic archive anyway. What, then, is the point of peer reviewed journals? Are we better off just posting papers publicly, and publicly commenting on them? Is stripping anonymity from reviewers a good or bad thing? Won't that discourage rigorous review, for fear of repercussions against reviewers? Is review really all that rigorous anyway?
My personal impression is that the volume of academic publishing has increased so much that it's impossible for readers to really keep up, and making it more difficult for scientific consensus to form completely. Publicly available papers in archives is a natural extension of this. What this means is that readers increasingly pick and choose which literature they read and cite, which truth they want to reinforce, and what truth they want to suppress. When you open up peer review to be public, those reviews become just another part of that literature. The peer reviews become blog and twitter posts, which is absolutely fine, but then a reader just selectively picks and chooses which what reviews and blog posts they cite, and so forth and so on.
For the record, I'm very much for open publishing, and open discussion of literature. I just think that academics hasn't wrestled with the implications of that, in terms of what it will look like (e.g., amplifying fads, decreased signal to noise ratio, increased feedback loops), and whether it's worthwhile to vigorously maintain anonymous peer review to have that as another form of literature evaluation. There's already a lot of public discussion of papers, and this will only increase regardless of what happens to peer review. Maybe the issue is who does the anonymous review? Maybe just opening up papers to anonymous commentary is the right way to go?
I've noticed this as well, and it just makes me think that quality reviews become more important over time.
Unfortunately, in my experience most reviews basically mirror what a couple recent reviews said, adding a few new papers. This is assumed to be up-to-date when in fact if the older reviews missed some important older papers, it's not up-to-date. And that's what I see: important papers missed by reviews in the past continue to be missed. I don't know if this experience is valid outside of fields other than my own, however. In my PhD I've tried to comprehensively review the literature and I've found quite a few important missed papers.
I think few people actively pick "which truth they want to reinforce, and what truth they want to suppress." My impression is that literature reviews are done more out of convenience than an intentional desire to distort the literature.
> My friends and I discussed just putting the commentary on an online archive, but it wasn't clear anyone would make the connection with the paper.
I think PubPeer is designed for this situation: https://pubpeer.com/
I don't think any of these biases are necessarily consciously enacted, but I think they exist in some of the ways you mention. There's a sort of echo chamber effect or positive feedback loop with citations.
Somewhere I remember reading a bibliometric analysis of citation patterns in the nutritional sciences regarding the effects of salt. There were two huge clusters of papers, one basically "salt is basically fine" and the other is "salt should basically be avoided". The papers in a cluster cited each other a lot, and not so much papers in the other clusters.
I agree that quality reviews become more important over time, but there's the issue of "according to who?" I suppose this becomes an expert judgment call and is the nature of these things, as it always has been, but I feel like as things become more parochial and balkanized the meaning of a "good review" changes somewhat, or becomes harder to agree on.
I don't assume that what you describe would not happen if the reviews were published as commentary; but I do imagine that the public nature of reviews would encourage more careful review, since in that particular case you mention the reviewer and the author(s) together would share the blame for the error.
Additionally, knowing what was reviewed would indicate what was not - and in your example, perhaps there was no careful review of the notation being used. That would be valuable for a critical reader to be aware of. I could read the article, see which parts were critiqued by the reviewers, and if I was skeptical, my critical analysis would build on the reviews instead of reinvent that wheel.
I'm not convinced review quality will improve with completely open reviews, because there's too many opportunities for retribution. Sure reviews that kill papers through a thousand tiny irrelevant cuts would probably go to the wayside more often, but my guess so to would very trenchant, highly critical reviews that make important but controversial points.
To me this makes the case for allowing commentary by non-reviewers pre-publication. The original reviewers may not find all the issues, and if someone happens to, by chance or because they're watching actively for things related to their expertise, there should be a venue for them to make a comment. I've seen a number of substantial back and forths for papers on openreview.net, often by people not initially tapped to review it, and found these discussions to often be as enlightening as the papers themselves, for instance https://openreview.net/forum?id=ry_WPG-A-
I've run into some of the same and it often reminds me of Kahneman's "theory induced blindness":
"we trust a theory so much that we search for reasoning to reinforce it even when a specific situation arises that our model doesn’t suit well, when we can’t think about other options, it becomes inconceivable that the model is incorrect or not suitable in a given context."
On the other hand, I've had items make it through the peer review filter that became obvious from the comments that the reviewers probably had little business reviewing that topic at all. It makes one rethink the value of the process.
I think it mostly helps improve the presentation of the paper by clarifying unclear parts.
Wakefield's "Vaccines cause Autism" study was peer reviewed. All the fancy powerposing/social priming/etc. studies that created the psychology replication crisis were peer reviewed. Bem's precognition research was peer reviewed.
As soon as you go into empirical research there are things that peer reviewers just can't check, e.g. whether the study author created the hypothesis before or after collecting the data. There's also things that could be checked, but usually aren't, like pretty much most of the software created for research.
It's poor phrasing, but in a review I would interpret this statement as "I made a strong claim, and my discussion wasn't strong enough to address the concerns of a reviewer who has only interacted with my work via this document." not as "you must be wrong because everyone else says something different". Reviews are a weird social process though.
Overall, you are absolutely correct that reviews are a really weird social context.
And it's very unlikely you've managed to upend a body of work you didn't even know existed, and has probably run into several errors before (which you probably didn't avoid, because you didn't look into it).
If you did read through the previous literature, it shouldn't be difficult to specify and compare.
It's possible for you to come from no background and upend everything.. but the chances are simply (and largely) in favor of a mistake being made. The onus is really on you to double-check and triple check your radical conclusions
Also there's just the fact that in CS we're working on a defined system. From top to bottom it's possible to definitively say when I execute X the following Y things happen. The details may be buried deep within the hardware for the truly eldritch bugs but fundamentally it's all human defined and the bits that do exist outside human design are well understood enough they can largely be ignored (no one is reasonably going to find a bug that traces down to some unknown quirk of semiconductor physics). In the sciences we're poking at the edges to see what the system actually is.
With computers we know exactly what we expect a piece of code to do and if it doesn't we're not discovering some new truth about the universe it's just that something was done incorrectly or our assumptions were wrong if it's someone else's code/program/library.
While it seems like a nice formal proof checker, it doesn't seem to have any automated their theorem proving abilities yet?
I can't imagine trying to do math where I have to manually supply a proof of every trivial statement... Something like Isabelle and it's sledgehammer automated proof finder seems more reasonable to me (which I hope to find time to dip my toes into this week).
Am I missing something?
Let's say we know that x = 3^a 5^b. It's a trivial fact at this point that x isn't divisible by 2 (by the fact that prime factorization is unique).
Disclaimer: I'm not actually good enough with lean to prove this sort of statement quickly... I could be under or even over selling lean here, hence why I phrased my original comment as a question about whether I was missing something.
If I was trying to prove this formally, I'd have to say something like (in computer speak): Suppose x is divisible by 2, x = 2 * y, y has a prime factorization p by <theorem in library>, and 2 * p is therefore a prime factorization of x (ouch, already not sure how to specify that statement). 2 * p = x = 3^a 5^b, so 2 * p = 3^a 5^b. Both are prime factorizations, prime factorizations are unique by <theorem>, therefore both expressions should have the same number of 2 terms, but the one on the left has at least one, and the one on the right has 0, so they don't. This is a contradiction, so x is not divisible by two.
You see why I'd like a computer to fill in the long formal proof instead of doing it myself? It's not because the statement might be false, it's because it's a pain to phrase everything in terms of the re-usable result about unique prime factorization.
While I'm disclaiming things, disclaimer 2: I'm a hobbyist, not a real mathematician.
I was formulating a vision with ergonomic tooling, helpful language constructs that haven't been invented yet, etc.
Something like lean isn't assembly, it's a full fledged language with abstractions, better IDE support than I've seen for any "real" programming language, the ability to program new "tactics" (methods of proof), and so on. I'm not trying to prove things in raw logic in it. The thing is I don't want to be trying to prove some of these things at all, I want the reader (the compiler) to just go ahead and supply their own proof.
Rittberg, Colin Jakob and Tanswell, Fenner Stanley and Van Bendegem, Jean Paul.
That is, the point is that the gaps cannot be trivially filled.