Best of Arxiv.org for AI, Machine Learning, and Deep Learning – January 2019
insidebigdata.com
insidebigdata.com
I get that ArXiv is a convenient way to download scientific papers, especially without access to universities or libraries' access, but one should always be careful with non-peer reviewed research.
but I agree with you that it's unfortunate that all papers aren't open. good things there's sci-hub.tw and libgen.io though :)
On balance it's good that more researchers these days will publish preprints before submitting for peer review. But there are still two specific dangers with relying on preprints:
1. Most people reading papers don't try to implement or test them. Unless there is a glaring error, it's hard to tell if a paper is critically incorrect without a lot of effort.
2. Even if most people did try to implement papers, that would still make for a poor heuristic on the paper's merit. In an ideal world every valid paper could be implemented. But even in conference proceedings, it's extremely common for papers to be missing details critical to their implementation. In many cases you can't do the implementation because it requires a vast amount of computation or proprietary data available only to the company whose researchers wrote the paper.
If the paper comes with code you can simply run to reproduce the results, that's a stronger signal for correctness that whether it went through peer review or not.
1. Most peer reviewed papers do not come with code.
2. Many that do come with code don't work out of the box.
the person i was responding to
>but one should always be careful with non-peer reviewed research.
what do i need to be careful about? developing the wrong intuition?
This is a dangerous line. Being a reviewer in CS, I don't think peer review is easy. The work in this field is complex and require a lot of time to correctly evaluate the claims.
First, CS/maths is not just code and theorems. There is indeed a lot of research which requires data and empirical validation, or propose new models.
Second, it's not easy to verify claims, even for code or theorem. This is why peer review requires many reviewers, often from 2 to 6, to evaluate one paper for a conference or journal. No one can evaluate in 5min whether a 40 pages paper on ArXiv makes sense or not. When I read a 70 pages complex crypto paper, I prefer a peer reviewed, community-backed work than a PDF not verified by other researchers. When you're a reviewer, you realize that many papers are often butchered, or contain (partially) false statement, not always easy to spot.
Anecdotally, I've occasionally received feedback in response to my posting a manuscript on arXiv which is as useful (or more) than what I've received from the peer review process for the same paper.
arxiv is where its at.