But actually, thinking back, 10 years ago there may have been more of a stigma around a paper that was only on arxiv.
But actually, thinking back, 10 years ago there may have been more of a stigma around a paper that was only on arxiv.
From the perspective of advancing science, journals seem mostly pointless and most of their benefit could be had at next to no cost if we just agreed to publish on the arXiv and pair it with a review system.
From the perspective of advancing scientific careers, because of how founders and institutions treat arXiv papers, you have to still publish in a journal, and so people keep publishing in journals.
But like, at least in physics time passes at the speed of arXiv posts, and getting published usually happens after it's been cited many times already. And if people are already incorporating it into the web of knowledge/responding to it etc., then what did the journal do?
And writing is thinking. Because you are forced to iterate on the text over a long period of time, your understanding of the topic improves and you learn to express the ideas better. Your job as a scientist is creating new knowledge, and the initial publication of new results is only the first step towards that.
Not sure if you understand how the sausage gets made. People won't spend effort at scale (there are always exceptions) on stuff that's not incentivized, such as blog posts. If you work on something, the goal is to make it count. But if you fail on that, arXiv is a second best option.
Also, papers that get published at conferences also get out of date pretty fast anyway. That's just how fast moving fields work.
But things may end up moving the way you are saying, as long as incentives keep up.
Journals (1-2 years) -> conferences (6 months) -> arxiv (weeks to write once the work is done, then days to publish) -> next thing (immediate AI writeup, code, results, tests, open scripts, weights, full reproducibility, full access allowed for other AIs to scrape and learn from, and reorganize and recombine on a speed of days or hours from result to adoption)
One of weird aspects of CS research is the tendency to publish small papers. In more experimental fields, a paper may represent years of work for the first author. But CS students often finish multiple projects every year. Maybe it's because of the focus on conferences. You want to attend conferences (to build your networks, and because it's interesting and valuable), but you often don't get reimbursed if you don't have anything to present. Students from other fields can just present their work in progress, but CS students have to come up with publishable results.
Maybe CS should also strive towards bigger projects. Instead of publishing individual results (which someone often improves a year or two later), the papers could focus on the outcomes of entire research directions. That way you are more likely to produce something of lasting value. Something that is worth writing up and rewriting and polishing, until you understand the ideas much better than after the initial write-up.
I wonder about the recent proportions. Growth in ML papers has been mind boggling and I wouldn't be surprised if a major part of Arxiv's submission spike was in AI ML topics.
> Maybe CS should also strive towards bigger projects.
The current rule of thumb is that you need at least three accepted papers in good venues to graduate with a PhD. Until that changes, the incentives point towards trying to publish often. And PIs want this because they also get evaluated in this publish or perish system. This is my main point. Throttling is a stopgap solution because academia doesn't want to admit its internal problems that have been boiling over already before AI.