Library-managed 'arXiv' spreads scientific advances rapidly and worldwide
ezramagazine.cornell.edu
ezramagazine.cornell.edu
You should upload your paper to arXiv. When you do, please upload your source (tex, or word I imagine), as well as a PDF.
For the blind, PDF is the worst possible format, and tex and word are the best formats. Don't hide, or lose, the blind-accessible version of your paper.
I've had arXiv automatically block bare uploads of TeX-derived PDFs (presumably identified through PDF metadata). For these, it required that the source be uploaded and compiled on their server.
On balance, it's probably best to require source as arXiv does, but this can create interesting issues from time to time. Since the source is downloadable, researchers can inadvertently end up sharing partial results or snark that was commented out in the TeX source.
I learned a lot about not underestimating people there.
It's a strange university but it always makes me happy to think that it (specifically Ginsparg, and to some degree strogatz and the law scholars) pushed forward what the Web was supposed to be.
Not a place to buy stuff, but a place to learn stuff, without waiting for it to make its way through a hyper politicized review process so that it could be printed in a 17th century fashion and mailed to some corners of the world, eventually perhaps reaching a fraction of the people who could use it. Rather, everyone everywhere with a connection.
Anyone who doesn't deposit preprints (arXiv, biorxiv, or wherever) or who doesn't agitate for their coauthors to do so is not really in it for the science. It's fine to be competitive -- deposit yours first. Make the fucking discovery instead of parasitically piggybacking on those who do the work.
But that last part, that's hard. Very hard. As long as there is enough money left in academia to encourage lazy shits, those of us who care about scholarship will have to push, hard, to remove the last refuge of these scoundrels.
You're either on the side of justice -- open data, open formats, open scholarship -- or you are tacitly endorsing Elsevier & Springer, who haven't the slightest problem using crap like incremental JavaScript & mangled PDFs to deny access to scholarship even to those who have paid.
David (blocking on his last name) at CTC made that choice for me. He set an example that forced me to admit what was right. I hope others will do the same. It's the right thing to do.
I think advocates of sharing science should also understand that, at present, the majority of the content is asinine or incorrect. Removing checks and balances is not helping.
It's not at all clear to me that the checks or balances were helping beforehand. Hiring a shit ton of MBA types to administer research probably didn't help much either.
Also, out of curiosity, is this open access / fast review problem what led to machine learning falling into irrelevance?
Because we wouldn't want that to happen to other fields. Can you imagine if raw clinical trial data was out there, for example? (The Horror! Best keep these appallingly expensive results locked up!)
Why? As the first author of several publications, more than half the reviews I received were not well thought out. This part of the process doesn't happen, Open Acess and self publishing websites exploited this fact for greater throughput.
>>machine learning
I would argue that the machine learning bloom in recent years is because of an availability of new technology from industry, supported by industry. This is also why they were able to tolerate less conventional venues for publication - as these were not viewed as mile markers for progress. Yet, when I look at more "settled" machine learning fields like speech, I see the same plethora of academic papers filled with useless crap.
Perhaps the growth of ML has nothing to do with Open Acess. There is although certainly there is an underlying profit motive for successful technologies, Google ain't going to use a technique that doesn't work.
>>raw clinical trial data
It would be great if it were published, and indeed many people would pay for quality trial data. I see OA and self-publishing of research results having no relation to authors withholding clinical data.
In the end it's the people who may or may not try to build on your work (or mine!) that are the best judges of whether it's BS.
Re: 2: probably, but diffusion of ideas benefits massively from readily accessible documentation. One of the things that will cause me to immediately reject a paper is if its method implementation does not work as described. I will always reject in such cases. They are alarmingly common and this is one of the few easy cases in reviewing. Lack of an implementation is also cause to reject.
Re: 3: it has everything to do with it. Unless you're running and publishing trials please don't lecture me on this. See opentrials for a particularly frightening take on why published trials so poorly represent those submitted to the FDA or registered with CTEP. Meanwhile actual RCTs get buried by shitbags demanding more subjects than there are recorded cases in the past 40 years (no joke).
If you're not doing trials, you may have to take my word for it (or get involved in RCTs and find out for yourself). A lot of the enterprise is actively harmful to science by omission.
If I just leave that out you seem to imply science publications should be selected by whether enough people care about, or whether they are boring or not. Both (especially the latter) are rather bad measuring sticks for the quality and importance of research.
Adding "financially viable" to "less boring" produces "EXCITING!", adding it to "enough people care" yields "POPULAR!". These are definitely not forces that should be pulling scientific research.
Not saying there shouldn't be checks. Just that market forces are probably too stupid for it.
Although that's not giving BuzzFeed enough credit, they tend to check their sources & verify results/reports. Also if a piece of fake news is simply resubmitted, they don't pretend like they never checked on it before.
There are plenty of checks (mostly from NIH, mostly for incremental garbage work on dead models like cell lines) and very few balances against the tyranny of CNS...
"If nobody cared about your opinion, it would not be financially viable to publish it. But now we have blogs/tweets."
Effectively, cost of publishing was working as a quality filter. Now, we need to actually implement proper filters instead of bad proxies for quality.
Would be nice if latex embedded the source of equations into pdfs. Wonder how hard that would be to add?
How large are latex files ? probably not much a few 100KBs. It's easy to add a compressed stream in a PDF. It's a great idea.
I'm surprised. Nobody has created an accessibility solution for PDFs after all these years of ubiquity? What's the story?
As an example, in post script the following example from Wikipedia would simply show the text Hello World:
%!PS
/Courier % name the desired font
20 selectfont % choose the size in points and establish
% the font as the current one
72 500 moveto % position the current point at
% coordinates 72, 500 (the origin is at the
% lower-left corner of the page)
(Hello world!) show % stroke the text in parentheses
showpage % print all on the page
Now if you're lucky you can just extract all quoted text and read those in order, but that's unlikely to work for all documents.It makes me wonder how screen readers work. Thinking out loud, it seems that the screen readers should let the applications (e.g., Word) handle their own parsing and presentation and obtain the data after that. Otherwise, the screen reader would have to reinvent many wheels, interpreting the code for all applications including all their versions, features, quirks, and platform integration issues - such a daunting and difficult task that it seems unlikely. But where do screen readers hook into the content? After it's output from the application but before it's an image for the screen (which could require OCR)? Ironically, I suppose PostScript or PDF could provide common interfaces.
PostScript and PDF have no such separation, the are the display logic. They are fully fledged programs that list the position, size, font, colour, of every symbol on every page, if you're lucky in a vaguely logical order, but there's no reason it should be. If the files were written by a human you may have some hope of extracting some of the content, but almost nobody writes PDF or PostScript by hand any-more.
With PDFs, all you have is the position of lines and characters on the page. There is no flow, ordering, or semantics.
It's not impossible, but I wouldn't know immediately what tools get this most right. And it's always a lossy operation going back and forth.
But this isn't something I've looked for in the past. So I wonder: in your experience, about what percentage of papers on the arXiv have included the source?
https://tex.stackexchange.com/questions/21838/replace-inputf...
The basic (and unfortunate in a xkcd/1479 way) issue appears to be that the arXiv compile server doesn't allow writing to subdirectories. A subdir \include implicitly requires write access for the aux file.
If one were to build a system along that line, meant to replace prestigious academic journals of today, of course it would be gamed. But isn't the general consensus that the current system already is being gamed and usually at expense of the researchers doing valuable research and the public at large?
(and in anycase, most subfields just have a couple thousand individuals involved at the PI level, who frequently interact at conferences, went to the same schools, shared an advisor, etc. So there doesn't really need to be an elaborate scheme to verify someones credentials. Chances are two individuals are already aware of eachothers reputations, or at least know some third party who is).
http://quantum-journal.org/announcing-quantum/ http://www.nature.com/news/open-journals-that-piggyback-on-a... https://www.aps.org/publications/apsnews/201602/arxiv.cfm
I review for others because others have done 1) for me. But I'll never review for Elsevier, and lately I've had the luxury of reviewing for the most cited of open journals (by operating bioRxiv, and accepting direct submissions from it, I claim that Genome Research is "close enough").
It makes me very happy that this is possible (my CV has not suffered for only publishing as first author, and whenever possible as senior or co-senior, in fully open journals). I'm pretty sure this wasn't possible for most people a few short years ago. That engenders optimism about the future of scholarship, for me at least.
Hopefully you as well.
On personal research, I've used it for exactly this, but since what I've seen was only preprints, I've often wondered about the final version. It looks like I'm not alone.[1] Do many or any of the arXiv papers get updates with the improvements that come from peer reviews? Is there a need for arXiv for finals or do publishers demand exclusives on finals?
[1] http://mathoverflow.net/questions/41141/should-i-not-cite-an...
Also, since you submit manuscripts to most journals in TeX, there's very little extra work involved in uploading the updated files also to arxiv. You maybe miss the copy editor's grammar corrections etc., but those are almost without exception unimportant --- also, more often than not, the copyediting by the publisher introduces errors not present in the original manuscript.
True, but I don't see that fights over precedence are unique to ArXiv either, or even made worse by it, no? I mean, at least now there is an unambiguous date-stamped public place to cite in this kind of fight. And those fights provide a built-in incentive to put stuff up there, which is good for all of us.
Basically: who cares about spitballs as long as the papers end up on ArXiv? Seems like a cost worth paying to me.
I'm not familiar enough with other methods of preprint publishing besides arXiv / eprint.iacr, but you may be right that it is not unique to arXiv.
My personal preference would to be to have bits of research done through something like git, so that work along the way can be seen, otherwise one may solve a problem and then be 'out-arXived' by someone who spends an all nighter tex-ing your solution (this is a hyperbolic example, but I think the idea of the potential flaw in the system should be clear).
I see what you are saying, but I don't think it's that cut and dry, otherwise I could just take someone else's work from yesterday (or whatever a short time frame is), and re-solve it (easily -since now the tricky parts have been revealed) and post it today - tada, I parallel invented it!
It's actually not that rare to have similar papers appear in arXiv one or two weeks later after you submit --- to me, it happened several times within last few years. In these cases, it is possible to see that the approach differs enough (and moreover, often you know the people in question, or you know someone who does).
Their point is that the stats are garbage-level useless. And I can imagine people bragging elsewhere that their paper received X,000 hits when in reality it's all spam or bots. It's not arxiv's responsibility to monitor that, but it wouldn't feel good to facilitate that kind of disinformation or invite hit inflation. Especially as scientists, we want to either publish good data or no data, not data that we know to be garbage.
It might be nice if ArXiv would perhaps provide the data to researchers on request. Just curious -- what kinds of questions would you use this data to answer?
It's not that I am personally concerned with misinterpreting the data. I just think there could be some downsides to releasing the data without limiting access in some way. For one, I think there are already issues with the citation metrics are used and interpreted, for example in tenure evaluations. I don't think it would be a step in the right direction if this data were used towards the same end...
On the other hand, perhaps a way for registered users to star papers that they like (similar to how Github lets you star projects) might be a good thing. It serves much the same purpose as a rough measure of popularity, but is entirely voluntary.
For that reason alone, arXiv is really helping the blind community in academia.
EDIT: Add missing 'not' :)
I think you may have mistyped this. ;-)
X is LATIN CAPITAL LETTER X
Χ is GREEK CAPITAL LETTER CHI
Everywhere I have ever seen it spelt, it has been written with the Latin letter. I have never seen it spelt as "arχiv" or "arΧiv", only ever as "arXiv".
I can understand why, for example, having a Greek letter in the URL would be undesirable, but if one is going to consider that letter to be Chi, then that should be the authoritative spelling and it should actually be spelt that way where possible.
Follow up question, how does a site like this have a $500k annual budget? I was napkin calculating the costs of running this and couldn't get anywhere close to $500k without having extensive staff salaries.
The other more greedy explanation is always money. Of course open source isn't antithetical to profit, but as mentioned before you do lose control and maybe Cornell doesn't want competition. Even if the project was started with the best of intentions, they still need to make it self-sufficient and maybe even profitable so they probably decided it's in their best interest. Of course this is all just me speculating.
http://onlinelibrary.wiley.com/doi/10.15252/embj.201695531/f...
Here's a discussion on HN of a blog post by me sparked by a conversation with Ginsparg.
My understanding is that they switched to the new domain after people noticed the original was being blocked as porn by a bunch of automatic content filters.
Long live open science!
Been pronouncing it "ar ziv" until now. :P
Machine learning is a field that's moving quickly and can also largely be checked quickly by other people in the field too. The problems that can't be checked easily are also those that won't be checked by the journals anyway (no journal I know of will retrain one of googles deep nets to see if they get the same result before publishing).
And please, I see what's happening here: I have been portrayed as the "traditional journal apologist" although it's definitely what I wanted here. I simply wanted to express my doubts that pre-printing is the answer to it all.
Anyway, is there a study that shows how many pre-printed articles in ML have been retracted or refuted so far? Until this is done, and shown as small as in peer-reviewed journals, keep a small basket.
I understand the process fully. It doesn't make it fast, nor does it make it a necessary cost to pay. Delaying access to content for several years does not solve a problem. A not-insignificant time was spent bouncing between people to sort out who was paying for the costs, then there is also a delay between acceptance and publication. This now averages just a month in pubmed, but papers can bounce around this point for a lot longer.
I am not arguing for pre-prints to replace traditional publishing, but the speed of spreading information is undeniably faster, and that's what started this whole chain of comments.
> Anyway, is there a study that shows how many pre-printed articles in ML have been retracted or refuted so far? Until this is done, and shown as small as in peer-reviewed journals, keep a small basket.
I don't get the phrase "keep a small basket", but no I've not seen this. The point is simply that PEERS (if we need to shout the word) can in many cases replicate the work and assess the results much more quickly than the traditional review & edit process. I picked the field because I see people re-implement work described in the papers very quickly, or the original authors share the trained models and code.
I'd also caution against using "retracted" papers as a measure, some journals charge for a retraction.
ArXiv is working very well for a lot of scientists in Machine Learning, since must results can be doublechecked by simply running the code. I won't disregard a proved approach that I've myself seen working just because it hasn't been published yet.
Could you provide some examples for a layman? To me it always seems that science moves far slower than one could hope for (e.g.: Every 'breakthrough' in batteries/Graphene of the last ten years and still no products)
Are you being literal about that or is that just a figure of speech?
Could you please provide a source for those claims, if true?
[0] http://journals.plos.org/plosmedicine/article?id=10.1371/jou...
[1] http://www.nature.com/news/1-500-scientists-lift-the-lid-on-...
Medicine on the other hand has the problem that you can't really afford large enough sample sizes to have sufficient statistical power.
This really isn't the major problem w/ modern medical research. In fact, if they had properly powered studies there would be far too many "discoveries" and the real problem would become obvious.
The real issue is that the efforts to come up with and study models capable of precise predictions (eg Armitage-Doll, SIR, Frank-Starling, Hodgekin-Huxly) have been all but choked out in favor of people testing the vague hypothesis "there is a correlation". There is always some effect/correlation in systems like the human body, so it is only a matter of sample size. As explained long ago by Paul Meehl, this is a 180 degree about-face from what was previously called the scientific method: http://www.fisme.science.uu.nl/staff/christianb/downloads/me...
Peer-reviewing has a lot of weak points, but saying that pre-print is the answer to them all is plainly wrong.
It's merely a supplement to peer-reviewed journals that has some nice characteristics, for some use-cases, which has been beneficial, to some researchers, in some fields.
What does "wrong" mean in this context? Does it mean that most scientific papers will be shown to be inaccurate in the due course of time? That is bound to happen given the nature of scientific progress.
But peer review aims to assess the methodology and rigour of the papers being presented. I agree that it is often debatable whether this happens in reality, but that doesn't explain your statement I quoted above.
I believe the PDF link is only removed after a paper has been withdrawn. But if you click on v1, you can still access the original paper (incorrect, in this case).
In any case I don't think that withdrawn papers support the OP's claim that PDF links are hard to find for normal papers.