Arxiv.org reaches a milestone and a reckoning
scientificamerican.com
scientificamerican.com
So what, you'll say, independent experts can catch your mistake if they read the paper on arxiv and contact you to suggest a correction. Yes, but there's no incentive to do the hard work to correct mistakes if your paper is already on arxiv and people are already citing it (who didn't do their due dilligence). Then errors start cascading. Science is self-correcting, sure, but it's much easier to avoid errors in the first place. That's an important function of modern peer review.
I'm speaking from experience, of course. Peer review has certainly helped improve my work. I didn't enjoy having to redo the work, but the end result was something I could be confident about. Also, I'm only speaking about publishing in journals, because peer review in conferences is a very different beast. In my field, of machine learning and artificial intelligence research, I'd go as far as to say that peer review in conferences is broken, getting published or not is a lottery and you're better off putting your stuff on arxiv: less hassle and more people will read it. You might even see your article posted on HN. When was the last time a paper published on a conference server was posted on HN? See?
https://www.youtube.com/watch?v=gv5mI6ClPGc
(Discussion w Noam Chomsky)
That's overly optimistic. What you mean is "with peer-review you know that someone will at least skim your work and maybe catch errors."
I agree with the rest of your points though. Zero peer review is probably not a great solution.
It's very possible to just keep submitting papers to journals until you get someone who shrugs and lets it through. Reviewers aren't paid, and while reviewing is encouraged it basically does bupkis to further your career (and academia is very cut-throat).
Is that a failure of peer-review though? Or is it more of a consequence that even if you do everything right you'll still get some false positives some of the time, and negative results tend not to get published at all?
(i.e. the problem that pre-registration for medical trials hopes to solve - https://en.wikipedia.org/wiki/Preregistration_(science) )
It raises questions around using peer-review as a mark of scientific quality. If it is, why does peer reviewed science time and time again turn out to have enormous glaring quality problems?
Peer review is not infallible, neither is it a mark of scientific quality, as you say. I appreciate that people tend to use it like that- you read, in articles in the lay press, that such-and-such study was "not published in a peer-reviewed journal" as if to say that we can't be sure of the quality of the study and that, conversely, if it was peer-reviewed, we could. That's not right. Peer review is not a sufficient condition to guarantee correct results. It's not even a necessary condition. Like I say in my comment above it is one mechanism of modern research practices that helps improve the quality of published papers.
Wrong, sorry. Disclaimer: I'm not a homeopath and don't go to any. Most of them are, in fact, quacks.
There are, unfortunately, chronic diseases which can be treated but not cured. Mainstream doctors and Big Pharma make a living off of those. If you have one of those, "well, why NOT try a homeopath?" is a perfectly rationale response.
"Sinus rinsing" is something that might be termed "homeopathy." It IS medically respectable, unlike most of their "treatments" (like magnets).
In fact, peer review is often more about novelty and importance rather than rigor. Very well structured research will get rejected if it isn't seen as contributing to the field in a nontrivial way.
A few fields (psych is the big one) are funding replication grants. That'd be the true mark of quality.
This is actually deeply problematic, as a part of what's driving the replication problems is specifically publication bias. If the only thing that gets published are unexpected or noteworthy results, you're selecting for statistical aberrations.
I think sometimes reviewers hide behind their anonymity to give snarky responses, and using their power over you, knowing you will try your very best to implement all of their suggestions in order to publish faster. I just think the whole process should be more transparent.
arxiv was about distribution. It didn't replace peer review - articles were still submitted to journals and published there too.
If an article was posted to arxiv and not a journal, the odds of a citation went down massively. And the journal it was submitted to was a factor in whether or not we read it. When articles were eventually published, most authors also updated the preprint with the post peer review version.
Basically it meant that (1) it was easy to keep up to date with what everyone was working on, and pick up interesting new stuff (2) most citations, post 80s, you saw in whatever paper you were reading, you could look up on arxiv and be reading it in seconds.
I'm surprised that they no longer use the term "preprint" at all, at least it's nowhere to be found on the homepage or "about" section.
The consequences of this amnesia are hilarious: https://twitter.com/gustavnilsonne/status/138948729731431219...
> Why do we call it "preprints"? The term seems to imply that work is preliminary or unfinished. As far as I can tell, the term introduced by @arxiv , the first online repository for scientific manuscripts, is "e-print". Is "preprint" a marketing device invented by publishers?
In machine learning, for the most part, arxiv is used to avoid peer-review. Or a way to "publish" work that has been rejected by a peer-reviewed publication, of course.
And to be more cynical, it's also a convenient source of references to pad up a Related Work section and make it look like incremental work is part of a growing body of groundbreaking new work. /jaded
Edit: well, I'm not just being cynical. The fact that everyone can put their half-baked papers on arxiv means that the 90% of work that is crap, per Sturgeon's Law, is now a much bigger quantity than ever before and one must sift through reams and reams of crap before finding work that has any meaningful results to report. Again, that's the case in machine learning specifically. I don't know about other fields.
Arxiv only lacks the initial quality filter by peer review.
I'm also working in the field of machine learning. In those niche fields I work more specifically (speech recognition), I can usually still get a lot out of Arxiv-only papers. I can pretty easily see the main idea and see if there is some usefulness in the paper or not w.r.t. my own research e.g. by good experimental analysis. In don't really feel overwhelmed in the amount of papers. I don't really see the problem.
In theory you should read arXiv and cite the conference version, but often people cite arXiv and nobody cares because google scholar mostly combines things properly.
Or, instead, you get a much larger audience reviewing your article, instead of the 2 to 4 reviewers involved in the paper publishing process :-)
You can get comments from critical and interested readers, you upload a new version of the paper to arXiv, and repeat the process until the article is ready for submission/publishing.
In other words: I think that an article on arXiv (on average and as a whole, individual articles may be exceptional) is subject to much wider and more extensive scrutiny and verification than the average article during the paper publication process.
So, yeah, I disagree. The level of scrutiny one gets from putting their work on arxiv doesn't compare with peer review by experts in one's field.
Discussion starts in blog posts, on Twitter, in conversations.
Then longer blog posts and code demos.
Then pre-prints and more blog posts about that.
Then finally journal.
You get incredibly less review at the publication stage — and your idea “experts in one’s field” don’t also use the internet is hilariously wrong.
I'm glad to hear you're amused by my comment. Where did I express the idea that '“experts in one’s field” don’t also use the internet'? Can you please show me?
I wonder if academic papers could benefit from a similar process of communal development (though it flies in the face of the dead-tree format of academic papers, but we're talking about Arxiv here, not Elsevier). Even if suggestions don't get merged, having a repository that captures discussion around the central artifact is also useful (such as when a maintainer decides that a feature doesn't belong in a project).
https://openreview.net/group?id=ICLR.cc/2022/Conference&refe...
Peer reviewing is essential, of course. But note that in most journals reviewers are not getting paid. They have even less incentives to find mistakes in your work than you are. I have no idea how such an utterly broken system emerged, where journals siphon billions of tax money exploiting the work of scientists.
This is quite separate than the exploitative aspects of the big scientific publishers. Reviewers are certainly exploited by publishers, because publishers make money from reviewers' free work. But reviewers would do the work for free anyway because that's what they expect others to do.
This is a remarkably naive (and charmingly optimistic) view of peer-review.
For the vast majority of the history of academic research peer review did not play a major part. It wasn't until the 1970s that the modern day peer review system emerged and it did so as a reaction to a funding crisis in academia to produce the illusion of legitimacy.
It of course has done no such thing. It has done nothing to limit the reproducibility crisis which has shown it's ugly head in nearly all scientific fields of study. One could even argue that peer review coupled with a publish or perish culture is the cause of such a failure of science.
This also means just about any great scientific discovery you can think of prior to 1970 did not go through the peer review process as we know it today. If science seems less exciting today, peer review is certainly one of the reasons for this. I'll leave you with Geoffrey Hinton's words on the subject:
> Now if you send in a paper that has a radically new idea, there's no chance in hell it will get accepted, because it's going to get some junior reviewer who doesn't understand it. Or it’s going to get a senior reviewer who's trying to review too many papers and doesn't understand it first time round and assumes it must be nonsense. Anything that makes the brain hurt is not going to get accepted. And I think that's really bad.
Laking peering review is a feature, and not just because it speeds up time to results.
Great way to join a convesation. Well done.
I didn't say anything about "the vast majority of history of academic research". I talked about my experience of peer review of my work. Are you responding to someone else's comment but quoting mine?
The situation for conferences in CS is basically how it works in journals in other fields. There may be more low-quality submissions to journals, but at the same time, the extreme competitiveness means most high-quality submissions are rejected also. The "solution" has been the appearance of open-access journals which are less concerned with subjective reasons for rejection like "impact", and scale to accept more papers. But they cost thousands of dollars. Open archives are a more democratic solution.
> But moderators frequently intervene, delaying posting by days or weeks, reclassifying papers or even outright rejecting submissions.
Definitely not my experience. I haven't posted to arXiv in a while now, but I wasn't even aware a moderation queue exists -- none of my papers were delayed, reclassified, or other...
EDIT: I see the article provides a few examples, but then they also provide numbers: "At arXiv, roughly 6 percent of submissions receive a hold, and about 2 percent are rejected." -- 8% of heavy-handed moderation intervention wouldn't qualify as "frequently" IMHO.
The arxiv moderators are like human spam filters. And as anyone who dealt with email should know, spam filters are essential for a platform to remain useful. I think the article puts undue emphasis on the few edge cases that invariably arise with any such filtering system.
And being a moderator is a thankless job indeed. The efforts of a moderator are largely focused on the absolute worst papers, and the job is therefore very different from being a referee or editor of a reasonable academic journal. So personally I would also focus elsewhere any efforts at increasing diversity in academic institutions.
It is peculiar that the model of fundraising via newspaper ads and a couple of wealthy donors (including in some cases the national government, yes) doesn’t seem to have survived past the world wars, but to say that private funding never resulted in any social good is preposterous. Something has happened in the last century or so, but it’s trickier than that.
Calling a hoogheemraadschap a private enterprise is calling a government one. If it acts like a government it is one even if it is not benefitting everyone equally (still no difference though).
That you can not rely on the rich to put their money in good use with the general public in mind is right. Does it happen? Yes. Does it happen often? Definitely not.
(similar thing in education involved deciding "equality" was bad, so they invented "equity", which is the same thing except they like it.)
> Other than some medical research, charity has never solved a single social ill.
This is very broad and incorrect statement that I suggest you research more before stating as fact. Maybe start with libraries in the 19th and 20th century if you don’t know where to start.
We still don't have a proper open /free tax filing too from the IRS or tax computed by IRS only for us to verify/confirm as in some of other countries, primarily because intuit (and others) lobbied to make sure the government would never develop one (encoded in law no less), so it is not that surprising that a community benefiting project that threatens some big business revenues has no government funding available.
It's incredibly annoying that publications try to continue to exist in order to act as a gatekeeper to exclusive papers. If journals added real value by peer-reviewing, collating, curating, offering editorial opinion etc then the fact papers are also available in the raw state somewhere else wouldn't much of a threat.
An online ecosystem like Arxiv can very easily supersede a journal if they want to.
For example Arxiv say could deploy a public peer review process i.e. reviews, criticisms and comments on the paper could be available next to paper in Arxiv without a lot of effort technically. Similarly curation and collation could easily taken up for free by community. Any of this would upend journals in quality or throughput easily. There are no problems journals solve that online cannot do better not even exclusivity they charge so much for.
After all there are enough people to moderate reddit or do various quality control tasks in StackOverflow for not much benefit, researchers who curate on Arxiv could easily be found and they may even get a job boost.
Building that community and acceptance is the hard part everything else is relatively trival.
--
Open Access especially for government funded research should be mandatory. Same for software, any thing government invests to develop should be open source.
If you're suggesting it as alternative for academic publishing in general, it would cost a lot more.
Arxiv isn't a "publication platform" in the sense that peer-reviewed journals are. The cost of hosting and distributing PDFs is close to zero, obviously. What does cost money is the copyediting and running the peer review process.
People think Elsevier are the worst people in the world (not entirely wrong) and taking in millions while exploiting the free labor of authors and reviewers. But it's relatively easy to quantify the value they do add: their financials are public and, last I checked, they have operating margins of about 30 %.
Let's say we could do without their marketing and billing departments and that gets us to a margin of 50 %. That's pretty good, but it isn't quite the rip-off people sometimes suggest it is. Replacing the commercial publishers with government-funded open-access journals would be a billion-dollar endeavor, not a few million for servers and a bit of software.
Running the reviewing process can also happen through https://openreview.net
Not sure why we need Elsevier exactly, could you be more specific?
Important to note that this is not some innate "want"; their careers depend on it. It's a means to an ends, the ends being being able to keep practicing science.
(Of course prestige-by-proxy will play a role as well, but it's not what's keeping the system in place.)
On the other hand, PubMed exists.
My guess is, for a preprint archive, network effects are paramount, and ArXiv was there first. It’s not directly a government project, but is still funded by US government grants and hosted by Stanford (now that the international mirror network is shut down, which does give me a slightly jittery feeling).
[1] https://hal.archives-ouvertes.fr/
[2] https://groups.google.com/g/isl-development/c/JGaMo2VUu_8
https://www.cadc-ccda.hia-iha.nrc-cnrc.gc.ca/en/
It probably helps that it is run under the auspices of the Herzberg Institute of Astrophysics.
Here in the UK, there is a big research evaluation exercise called the REF. The REF is important because it largely determines how the government allocates funds. In order to be eligible to be evaluated, researchers have to essentially submit their pre-prints into open access.
What this has meant for the last 10+ years is that most universities participating in the REF have an e-prints server, like the arxiv. Obviously, being publically funded, universities are then using government money to pay for these costs.
Now your question might be whether it's sensible for the entire country (or research councils) to operate their own e-prints server. I don't know the answer to this.
All I know is that the more bureaucracy you involve, the more bullshit results.
So on one end you could takes something like Newtons Principa, that synthesised a lot of the known facts and laws at the time into one single coherent framework, and on the other you can have Galileo recognising that he was seeing surface features on the moon and not just patterns on a flat disk and announcing this will some illustrations showing the principle and not what he actually saw.
As to the first part of the question, if you talk of "publishing" in academia you usually mean you want your work to be know by other credible researchers in the hope they find it useful.
The most reliable way to achieve that without credentials is to befriend someone respected in the field you are interested in before you send them your results, because they are more likely to have patience with any errors or confusion in your presentation that way.
There are of course some places that pride themselves on being completely free from filtering, but the concept of "spam folder" along should tell you how good a chance you have of getting any serious attention in those places.
A more approachable avenue could be blogposts and videos. One good example of a scientific blogpost I recall is here: ( https://dynomight.net/2020/12/15/some-real-data-on-a-DIY-box... ) Someone tested very simple air purifiers with limited experiments. It follows the same "abstract-methods-results-discussion-conclusion" format you might expect. Some obvious caveats, this data could have all been faked, and it isn't peer reviewed for methodological flaws. But you can reproduce the experiments yourself for ~$150.
I am in academia but I like blogposts a lot more. I wish they were more popular. I think HTML makes more sense than PDF, and to some degree they cut the middleman out between communicating to peers and communicating to other academics.
HN is a nice tidy list of top things, is there an equivalent for ArXiv?
I just have a few bookmarked searches for what I am interested in.
I don't see how you could generalize things though across all papers. Like randomly clicking on what was submitted to High Energy Physics - Lattice yesterday. I don't have the first clue what any of that is and that would be the same for most topics for most people I imagine.
My opinion here, but it sounds like instead of a 92% immediate acceptance rate, the author would prefer an immediate acceptance rate closer to 100%. I’m not sure why diversity would get them there, but I am encouraged that the current moderators are able to produce 92% immediate acceptance rate. On the other hand, I am distraught by the suggestion that a highly diverse moderator base would presumably only account for a few percent reduction in holds and rejections. It’s my understanding that that should have a far more significant impact.