Arxiv Sanity Preserver
arxiv-sanity.com
arxiv-sanity.com
Karpathy's Arxiv-sanity helps a lot to keep in touch with the latest and greatest deep learning without having to spend all my time reading papers.
What are all those papers like? I wonder if a large proportion are the equivalent of peer-reviewed blog posts.
[1] https://www.quora.com/How-many-academic-papers-are-published...
To combat this a fair number of people will publish many times on the same developing experiment (not that this is necessarily a bad thing).
Don't forget grad students and dedicated undergrads. That adds a fair bit to your existing numbers.
The easy way to answer this would be to use something like Web of Knowledge to get a rough sense of the number of distinct authors.
Whatever that means.
That's a first use case. The second way things are out of hand is that you remember this paper from 3 years ago that was very related to this one, but can't remember it's name anymore. Here you can sort by similarity to any paper, and usually these papers come up on top of the sorted list. This is also useful for finding related work. Another use case is a peace of mind that you somehow did not miss some papers that you definitely should know about.
Google Scholar is supposed to have similar features: it emails you papers it thinks you would be interested in and can in principle show similar papers. I don't know what they do internally but these features are quite terrible and low quality in my own experience compared to what I get here. More generally the amount of innovation in Google Scholar over the last few years is sadly either zero or negative (but overall I still get nightmares about what would happen to academia if Google pulled a Google Reader with Scholar). For arxiv-sanity it's tfidf vectors of bigrams from full text of each paper and I do L2 lookups for similarity ranking and train personalized SVMs for people for recommendations. The results are, at least for me, significantly better.
I've been reading papers from the 1960s, which is when the term "information explosion" was coined. People then were struggling to stay current with the literature, and thought 'things were seriously getting out of hand.'
This was the start of abstracting services, like ISI, where you could even arrange the results of a keyword search of all the new papers to be sent to you each week - a clear predecessor to personalized RSS feeds.
Going back even further to the immediate post-war era, the library systems of the time, which were structured around books and journals and organized by topic, couldn't keep up with the deluge of research reports which cut across multiple topics. The field of information retrieval, using first punched cards and then computers, started because the publication flow was 'seriously getting out of hand'.
Or for a specific example, after high T_c superconductors were discovered in 1986, there was a mad rush of interest as solid state physicists from around the world explored the new territory. A Google Scholar search for "high temperature superconductor" finds:
1986 - 846 publications
1987 - 2 600
1988 - 3 900
1989 - 4 780
1990 - 4 870
1991 - 5 250
That's 14 papers per day, any one of which might be "very related to your research, or scoop your latest idea, or have good ideas you can use in your own work."Granted, 14 << 50, but that doesn't include some of the papers about "high Tc" which don't use the whole phrase. Also, those are 14 peer-reviewed papers per day, so there has been some filter, and experimental research in high Tc research requires more equipment than deep learning.
Think of my comment as a reminder that things have been out of hand for most of a century, and dealing with that deluge emotionally connects you to the headache that generations of researchers before you have had to suffer with. :)
Given how low the reproduction rate in science is, I'm not sure that time would be wasted.
For example, suppose you discover the structure of DNA, you try to publish but find someone already published it last year. You've just wasted a lot of time.
I don't know what the solution is.
In the fields of computer science where you actually implement something, it's the design and implementation of that new artifact that takes up most of your time. The experiments are not nothing, but they tend to be the sort of thing you can script: run a bunch of programs, accumulate results, do data analysis. Even the data analysis tends to be scripted. If, at the last moment you discover a small tweak that could improve your implementation, it can be trivial (in effort, not necessarily time) to re-run your experiments.
Now, the experimental evaluation is still important, and it is also easy to do wrong. But I also claim it's more deterministic. If an author is honest in describing their experiment, it's easier for reviewers to cry foul in computer science systems research than in, say, psychology. There are ways in computer science to design poor experiments that show bogus results, but it tends to be more obvious.
If, upon trying to publish, you discover that someone else had a similar idea and implemented something similar, you have replicated the hardest part. You spent a lot of time and effort designing this new thing that overcomes all of these challenges. If someone else already did that, you could have just skipped all of that, and started on improving it right away. In computer science systems research, replicating someone's research may actually be must faster. Sometimes you can view their code directly, or you can implement their idea in another system. Re-implementing an idea in a new context can take a lot of engineering effort, but it can still be a lot less work than doing it the first time.
Now, what happens if that was a bogus technique, and you can't replicate the results? That's a publishable result, but it tends not to be the whole paper. You figure out a better way, and explicitly compare your new way to that old published way. Again, that's because in computer science systems research, you're not discovering fundamental properties of things. Instead, you're discovering better ways of doing things.
I do sometimes read computer science systems papers and think "Eh, I don't buy this result". That's usually not because they did anything wrong (although sometimes it is), but because I just think that what they are investigating does not matter. "Sure, I believe you figured out a reliable way to optimize a three wheeled car, but four wheels is still better."
Theoretical computer science is not impacted by this at all, as there usually are not experiments in such papers. Their "results" tends to be a proof.
Others in this thread propose RSS and aggregation of abstracts as a solution. My proposal is to have reviews of literature, and then just read the reviews instead. This should save a bunch of time.
Thank you Andrej for putting this together and maintaining it.
I hope it will support TLS for registration or login mechanisms.
One of the problems I caught: error: [Errno 24] Too many open files from tornado. Trying to fix (edit ok made ulimit -n larger and I don't see this error anymore at least)
Does this happen always with the same traceback (that is, with the same area of your code in the stack trace)?