A Reboot of the Legendary Physics Site ArXiv Could Shape Open Science
wired.com
wired.com
The HTML is small and loads fast and automatically works on mobile because it's just simple HTML.
I hope the reboot won't mean huge bloated HTML+JS, and then a "mobile" version that has less features than what it has right now.
P.S. the wired article talks about arxiv.org but doesn't have a single clickable link that goes to it??!!
> As one of the first open access science sites, this redesign may influence the paths of its younger siblings, like biologists’ bioRxiv whether they follow the same route or rebel against it.
There's a link in the article.. which links to a Wired article on bioRxiv[0], but not to bioRxiv[1] itself. And before you ask: no, the article on bioRxiv does not link to bioRxiv.
[0] http://www.wired.com/2016/02/the-rainbow-unicorn-trying-real...
They could've at least added a useful outbound link with rel="nofollow" to be helpful while avoiding leaking any of their precious juice.
As I see it, it would solve many issues. Ideally, we would move away from a publication-count metric, and more onto a reputation-based metric. It would lower the bar for participation in scientific discussion, make reproducibility more important, and generally be a healthy thing for science as a whole, I think. But I want to hear opposing viewpoints!
Although citations are horrible metrics, I think some sort of reputation ratings based on quick likes would be even worse and even easier to game.
But there is the rub: if you make comments open, then you will will probably be flooded with. If you close them up, then fail to get the benefits of openness.
That's the thing, though - speaking as a former academic, the reaction when you see a paper that needs serious commenting is mostly "I'll just ignore this paper, since it has these issues", and you continue working on your own stuff - which again is because of the need for citations. If comments had larger effects on your career, they would be worthwhile making. I would guess the effect would be less papers and more 'collaborative science' where you start from a paper and either suggest improvements or improve it / replicate it yourself. I don't see the current proliferation of papers to be a good thing.
As for the worry that there will be too many 'laymen' doing the commenting, there could (as suggested elsewhere) be a verification process - i.e. there could be a 'Verified Ph.D.' (or some other measure of merit) comment section in addition to the 'layman' section.
There could also be a 'replication' section, where replication efforts will be rewarded - possibly by giving the replications some fraction of the paper's total 'reputation'.
In short: If comments are made to really count for scientists, I think there will be a more healthy scientific process going on. Viewing a paper as an evolving thing is IMO more in line with how science does work anyway - a result should be replicated before it's accepted, which is not the case now.
Let's not fool ourselves, not all refereed papers are great, not all of them are significant and not all of them are published in high impact journals. Most of the reviewing is cursory, menial, and often delegated to postdocs/phds. Most scientific topics are highly specialized, and you 're unlikely to see trolls bothering to comment on them. Writing another paper as response is not a solution either: the pace of article publishing is months and years, not seconds.
An example: I have often found the discussion of scientific papers here, in HN, to be illuminating, clarifying or countering issues, and offering a wider perspective that is often not mentioned in the article itself. And it's not like everyone in HN is a luminary, just mostly inquisitive people. I also see articles discussed in a useful way in twitter. Why can't we open up this discussion?
I don't think comments are to be taken as reputation metrics either. But i do think a simple, helpful open discussion section is missing from every paper that i have read. I believe the main barrier to it is that academics do not want to get off their high horse of untouchability.
To be precise, many journals do have a comment section, but nobody uses it.
GitXiv[1] also has a comment section and voting features, and nobody uses it.
I know it has a small user base compared to ArXiV, but I think that its format resembles what ArXiv might look like with comments, votes and reproducible code.
[1]: http://gitxiv.com/
Likewise SciRate is a reasonably popular arXiv overlay, but the comments are never used: https://scirate.com/
For example, a journal might have a "Letters to the Editor" section for people to voice a serious comment, even though such a letter is not "another paper".
As a fictitious example of what could go wrong. I am a theoretical astrophysicist. I am not established (I am not a tenured professor). Suppose that I come up with a good model to explain some astronomical observations. What I would want to do is to write a paper directed to my astronomers colleagues explaining how my model works and why it is better than competing models. I want to convince them to work with me to analyze their data. As it is now, I would post it on the arXiv, the astronomers would probably find it, read it and evaluate it. However, if the arXiv had comments, a single negative review by a more established theoretical astrophysicist would be enough to discourage any astronomer from even reading my paper. Remember that in the astro-ph section of the arXiv there are of the order of 100 new papers per day and we can realistically read only 1 or 2 papers per day on average. In this situation, the chances of my work being completely ignored because of that one comment would be very significant.
I think that the current channels for commenting on scientific work, private email and/or rebuttal papers, are perfectly adequate.
With due respect, do you have experience inside the world of theoretical astrophysics research or something similar? Do you know how it actually works or are you speaking ... theoretically (ahem)?
It is also true that a famous professor can backstab you when reviewing your paper. But, first of all, your work is already on the arXiv, so everybody already had a chance to form their own opinion. Secondly, in a peer review process the reviewer cannot arbitrarily reject papers. It does not work like that. He/she has to provide good motivations for his/her recommendations. Reviewers can also be challenged to the editors, who can ask for second or third opinions.
Finally, as others commented, to truly disseminate your work you need to go out and engage the community, giving seminars and talks. I couldn't agree more with this. My issue with that, and I talk as a privileged because I work in one of the top Universities in the USA, is that going to conferences, giving seminars and so on, is way easier if you come from one of the top places. You have funds for traveling, you had a lot of chances to network with the right people (when they visited your institution, for example) and so on. It is much more difficult if you come from lesser known groups or from abroad. The system as it is now already strongly favors people working in the top research Universities in the USA and Canada.
The arXiv is great because it puts everybody at the same level. It ensures that the best ideas have a chance to come out, independently from their origin. I wouldn't want the scientific discussion to be dominated by few loud voices.
With arXiv as a basis for comparison, what is your impression of what the other / older methods of dissemination overvalue and undervalue? For example, my impression is that the other methods favor top U.S. research institutions, but maybe that's realistic; maybe the older methods actually undervalue top institutions (despite my egalitarian fantasies). Maybe gender or experience or position or scope or novelty or other things are over/undervalued.
I do notice that in scientific research, institution is almost a surname in people's identities. It's always 'Jane Doe of Harvard'; it seems like it might as well be 'Jane Doe Harvard'.
I fully think open peer review in various forms, is awesome and should be pursued, but the arxiv serves a more fundamental need as well: To simply make the papers available. Within the fields where arxiv is established this is not an issue, but we need to get adjacent fields, and really everyone to agree to a publicly available database of research first. There is still considerable reluctance in some quarters to upload to the arxiv as I found out after changing fields.
The arxiv should be a pure database, then we can think about various ways to build on top. That could be arxiv overlay journals or journals that try out more fancy, new peer reviewing systems that might include curated comments [0].
As long as there is considerable disagreement over what the right format for serious commenting on papers is, arxiv should not be taking sides.
[0] http://www.earth-system-dynamics.net/peer_review/interactive...
The only danger is that the true obnoxious nature of some academics would be aired in public. There is a lot to gain from opening up a discussion in each paper. Asking questions, making clarifications, even suggesting improvements are things that are not possible to do now.
I'm not sure how you can say that. Generally, the comments on the vast majority of internet sites are pretty low quality, and even insightful and useful ones are difficult to find because of all the noise.
Because others can extremely negative - to the point where they may put some people off publishing in open forums.
The very small amount of value that they may add is easy outweighted by a single person not publishing.
However, some general questions being publicly available would be also arguably have utility.
i actually agree with this
though i do think comments and peer review would be great for arxiv papers.. i think it should be done by someone as a webwrapper
No, I think the major feature about having the arXiv host comments directly is that it would force physicists to engage with folks who comment on their paper, because otherwise there would be unanswered criticism attached ("stapled") directly to their public-facing papers. Maybe that's what we want, but that's a huge step and there are a lot of ways for it to go badly.
Here is a mockup that Paul Gisparg made for author-curated links on arXiv article pages. See the "Author suggestions" column on the right: http://www.cs.cornell.edu/~ginsparg/arxiv/1212.3061-mockup.h... Allowing links to places like SciRate plausible solves most of the chicken-or-egg problem by allowing the author to designate "the" place for public discussion without picking a particular site to win or lose, or bundling release of a paper with forced public discussion.
Shameless plug: folks here may be interested in my blog post about the future of academic papers and how the arXiv influences that: http://blog.jessriedel.com/2015/04/16/beyond-papers-gitwikxi... There was good discussion on HN last year when it was posted: https://news.ycombinator.com/item?id=9415985
Here's my 10-minute take on scirate (from viewing it on a mobile browser). Overall, I don't immediately see any added value. The layout and design isn't great - not on mobile at least. The main view isn't optimized for what I want to see. There's a huge navigation tree that takes up all the view space. Then, it shows me a list of ranked papers that are not of interest to me. It forces me to register to access any features. I don't see comments on any of the papers, only a Reddit-style rating. Do I have to register to see the comments? Overall, I don't think I'd ever go back. But, maybe it wasn't designed with my type of use cases in mind.
Personally, I don't expect it to be nice on mobile and I don't want it to be effortless to comment. The amount of work it takes to make a useful academic comment dwarfs those tiny frictions.
Also, ArXiv has a lot more than physics. Here's a map over all the papers: http://paperscape.org/
It might be a bug though.
They estimate spending just over a million dollars this year, funded via members (university libraries contributing a few thousand dollars each) and grants.
However, as I always point out, we must remember that the costs of a published journal are higher for reasons other than profit: journal articles are typically converted to HTML and XML via semi-automated processes with manual proofing and tweaking, production staff and some editors (for the biggest journals) are paid salaries, peer review involves chasing down deadline-abusing academics over weeks or months and building some kind of review tracking system, articles have to be reliably archived (e.g. lockss.org), and so on.
Then add that most journals are still printed and distributed by mail for some reason, and you have extra full-time staff and equipment.
I wouldn't want to throw all of that away and just keep the arXiv. Ditching print makes sense, some things can be better automated (we need better tools than LaTeX or Word to prepare papers, so they can easily be produced in PDF, HTML, or semantic XML form), and some services are unnecessary (I've had journals copy-edit my articles and make no useful changes), but I don't want to just read crappily-formatted base LaTeX from the arXiv with no frills either.
(Incidentally, PLOS is non-profit and still has to charge over $1000 per article in PLOS ONE, to pay for the professional staff and all the manual manuscript conversion gunk. Biologists love Word, so converting must be hell.)
What would you suggest as an alternative to LaTeX or Word?
I'm just curious; I use LaTeX regularly and think it works just fine.
I don't know of any good alternatives. LaTeX succeeded because of its excellent extensibility, so the core can be used to produce almost anything. I'd like to find similarly extensible tools that can cleanly produce XML as well as PDF. Maybe Pollen (http://docs.racket-lang.org/pollen/) will be one.
There's also Scholarly Markdown (http://scholarlymarkdown.com/), but I suspect Markdown's lack of easy extensibility will doom it, since you'd have to write documents that fit narrowly into the features they provide. (What if you want a "remark" environment and they don't provide it?) reStructuredText is also an option, since it's built for extension, but it's not very well-known outside of Python circles.
The problem is that most peoples frame of reference when talking about comments and "likes" is the facebook, reddit or [insert some publisher] model. That is all fine and dandy as a starting point, but let's not forget that all of those models are specifically engineered to keep people posting and liking as much as possible rather than focusing on like and comment quality. There is no reason, or at least i don't see it, why a different model focusing on like and comment quality could not be engineered. Furthermore, if such a model could be engineered well enough, I don't see why a comment and like section of arxiv could not contribute to furthering science as a whole. With that in mind a much more fruitful and positive discussion could be had about "how do we create a comment and like section where everyones interests are protected?" rather than the old back and forth discussion of easing barrier of entry vs quality.
As others have noted much better starting points for the above discussion could be the stack-overflow model or something similar.
Arxiv could use some of that reputation & moderation structure to regulate content. Cranks and time-wasters could be sidelined. It would certainly work much better than a typical newspaper comments section.
It would work by establishing expectations of how much substance should be in a response to a paper. Then getting the community to pitch in on the quality of responses. You still get echo-chamber effects with this, but the quality filter is generally worth it.
Groups of academics could band together to form online-only journals on a specific topic and then allocate medals to papers that are considered accepted. So while there may be many loose papers out there with little or no comments (like on axiv currently) only a few will get the seal of approval.
https://twitter.com/search?q=%23LinkedReproducibility
https://twitter.com/search?q=%23MetaResearch
- schema.org/MedicalTrialDesign enumerations could/should be extended to all of science (and then added to all of these PDFs without structured edge types like e.g. {intendedToReproduce, seemsToReproduce} (which then have specific ensuing discussions))
- http://health-lifesci.schema.org/MedicalTrialDesign
- there should be a way to evaluate controls in a structured, blinded, meta-analytic way
- PDF is pretty, but does not support RDFa (because this is a graph)
... notes here: https://wrdrd.com/docs/consulting/data-science#linked-reprod...
(edit) please feel free to implement any of these ideas (e.g. CC0)
a way to evaluate premises (assumptions, controls, data, data transformations) and conclusions (presented as e.g. JSONLD, RDFa in a standard form (potentially like IPython/Jupyter .ipynb; but with an OrderedMap of I/O sequences with fixed #urifragment IDs)
Sorry Wired, but too many newspaper websites have abused their advertising privilege for me to ever read a newspaper without an adblocker again.
I kind of like that model. Have the site being good at what it is best at, and help other services provide further services, such as search (google scholar) or stricter moderation and commenting and discussion.
> 1: of, relating to, or characteristic of legend or a legend
> 2: well-known, famous
It seems you only knew of the first definition, when the article uses the second.
I have the no-tracking enabled in the FF preferences and I'm on Ubuntu MATE 16.04 LTS.