Do I really have to cite an arXiv paper?
approximatelycorrect.com
approximatelycorrect.com
edit: To be clear, I am personally entirely okay with not citing papers containing ideas you were unaware of during your own formulation (though I think if you become aware of it, you should probably point out that it was previously independently discovered by someone else). Your paper may still have merit even if it's idea isn't "new" (especially if the first paper is shit as is often the case).
edit2: I personally don't see much wrong with Yoav Goldberg's blog post linked in this blog post. It's refreshing to hear his honest opinions out loud. As a graduate student I always lost a little sanity each time I read a paper with "great ideas", but terrible follow-through (i.e. explanation and proof of those ideas). I personally think that clarity of exposition is at least as (if not more) important than the novelty of an idea. However, you should still cite the sources of your ideas. Feel free to point out the source's flaws, but cite them nonetheless.
(Hopefully no more edits...)
What I'm saying here is that the "academic code of conduct" is a bit outdated.
Of course tangential to the citing of novel ideas is the citing of material that helps explain your work. I think it's certainly a good idea to cite a source of common knowledge if that source does an especially good job of presenting that knowledge.
As with all things, this is a judgment call. Just try to be honest...
"To solve this discretized Poisson equation we use the BiCGStab method \cite{vanLeersPaperAboutBiCGStab}, which is an iterative Krylov method; see \cite{SaadsTextbook} for a general introduction."
the bibliography is an important part of your work. its there to help an interested reader understand your work more fully by placing it in context, providing more access to a more exhaustive discussion of the finer points, and to assist in their studies of related topics.
at some level thats what it has to be about.
Sometimes you also cite sources to help the reader. E.g., I have cited Wikipedia articles for basic concepts in peer-reviewed articles.
1. Assigning credit to where you heard about something.
2. Giving a claim (the best possible, and sufficiently strong) support.
From the credit perspective you should cite wherever you heard something, even if it was alleyway graffiti. From the support perspective, you should take the idea you read on the wall look it up to see if someone credible has said the same thing, and then cite them, and if not maybe not mention it all.Since publishers are commercial entities, their prestige is an asset they want to protect.
So there are other incentives at play when it comes why citing ArXiv is currently not en vogue.
I always find it refreshing to see how pre-print driven reserach communities like physics operate in comparison.
"If you haven't climbed the ivory tower you don't get to speak in the ivory tower... Also, we in the ivory tower don't listen to those not in the ivory tower. Less we ourselves are cast down from our perched position."
3. Show what is novel in your work. Novelty is important in many academic situations.
However, sometimes it happens that you read first about a topic on wikipedia, and worked something out on the information you found there, before you had time to consult wikipedia's sources. In this case you definitely should cite wikipedia.
I was advised to always cite the primary source, even if I learnt of something in a secondary source, eg. lit review or Wikipedia. The reason being that if the reader wants to follow up on it, it's the quickest path. I'd say it's also just good practice to credit the original authors for their work.
Citing Wikipedia also puts the reader in the uncomfortable position of either taking your word for something, or having to go through Wikipedia's sources themselves to verify something.
How does this differ from citing anything else?
https://www5.in.tum.de/~huckle/Sutton_Spinach_Iron_and_Popey...
I must add that this was such a pain for me, as I found several relevant articles on arXiv and I could download and read them. I can't say the same for articles found on Elsevier or ACM, where the relevant articles were mostly in the journals to which my University did not have access...
Citations to a finalized version of a published, peer-reviewed article are "best", both in terms of assigning credit (this is what the authors are supposed to be producing) and as a pointer to more information for the reader (the article has been reviewed[0], it won't change, and there's a stable location for it). Work that isn't peer reviewed shouldn't be outright banned or ignored, but the citation should carry a lot less weight. It hasn't been reviewed, it's subject to change, etc. Since these are essentially someone's musings on a topic, when you cite paper to "prove something" (e.g., you write "The work of XYZ et al. (2017) shows that <some confound> is not a problem"), people will give it correspondingly less weight.
There is a long tradition, predating arXiv by decades, of citing technical reports or "white papers". These are usually written up like a journal article, but might be difficult to publish (all negative results) or contain more details than a typical journal publication would allow. If there is a "journal" version and a "tech report" version, it would probably be better to cite journal version, but I would be shocked if someone actively objected to including a tech report.
(In some disciples, the white papers are also the only thing available. The World Bank and Federal Reserve, for example, often release white papers containing their own data. They rarely bother to publish them in a journal though).
[0] For whatever peer review's worth :-/
I couldn't resist writing that so I read the article looking for something substantive to include as a sop to my conscience.
He didn't contextualize or establish the existence of the problem to my satisfaction. A naive reader would assume that he is that he is arguing against academics who feel literally entitled to plagiarize pre-print publications. I don't buy it. I suspect that the actual debate he's engaged in involves is better characterized by the following three quotes:
"…many authors are peeved, pricked, piqued, and provoked by requests from reviewers that they cite papers which are only published on the arXiv preprint"
"Any time that our work follows … ideas from other people, and when we can reasonably be expected to be aware of this, we ought to cite the related work."
"If similar work comes to our attention during a proper literature review, we ought to cite it."
To which the counter-argument would the following, in his closing passage:
"Many reviewers are abusing the system and asking for ridiculous comparison to recently-posted preprint papers…"
This acknowledgement comes too late to be given any useful answer, even though one can easily see it to be the core of a real problem; the reviewers, after all, are serving as gatekeepers to publication, and if one should not put "too much faith in … the overworked cohort of peer reviewers, roughly 30% of whom typically fail to even comprehend the basic outline of the paper," then it is probably extremely frustrating when someone insists that you cite papers that you haven't read in your bibliography as 'related literature', let alone papers you have read and dismissed as insufficiently important to refer your reader to, let alone papers that were clearly written by the reviewer's pet pony in crayon on a stable wall before being photographed and uploaded to the arXiv as uncompressed IMG files.
I guess it's a good reminder that many of these habits aren't purposeful. Though I'd hope that people would stop and think a bit more than I just did before uploading papers to the archive (or publishing elsewhere). But I guess impatience afflicts us all...
(This time not an edit!)
Im not saying it's necessarily a good system (plenty of shit makes it in to journals after all), but I can understand why an author would be hesitant to cite lots of non-peer reviewed sources in a paper of theirs. Having said that, I guess that's really just a sign that you're on shakey ground if you're reliant on dodgy sources!
I think no matter the source, you (as the author) have the final responsibility of citing correct work. You can (reasonably) choose to only cite nature because it is "safer" choice, but you can also cite other sources though you should of course take care to vet that source well yourself. As you point out, a lot of shit makes it into journals and in fact many journals are themselves shit (and outright frauds) so adding arxiv as a source doesn't really fundamentally change anything.
As a side note, I believe that many papers cite way _too_ many papers. Unless you use or expand upon a paper's work, I think you shouldn't be citing it (except possibly as general background knowledge). I just can't understand how you can write a paper that's explicitly doing so with 100 previous works (a book sure, but not a paper). Then again, my citation philosophy goes against many others (including my former advisor). Many think you should cite basically anything tangentially related (especially the works of the academic king makers). For me this is no longer important since I no longer am an academic researcher. I have the freedom to pontificate on the subject without any worries of having a career. :)
It allows an author to state the reason for the citation (supports, refutes, etc.).
Having to state a reason could reduce spurious citations while providing a more accurate view of the nature of a paper's impact.
In fact, I personally lean towards often including arguments even if they can be found in sources you cite. If you can provide a much clearer argument than the source (and it doesn't detract from your own work), you should include the improved one in your work. If it is a minor detail that is both worth citing as well as easy to include, then I think you should include it to avoid requiring the reader to hop around from paper to paper to get an understanding of your work. You should make your work as accessible to the reader as possible.
If I'm writing writing a paper about my ideas, it would be unethical to "make my friend a co-author" if my friend didn't contribute to those ideas.
Yes, it would be strongly preferable to tell readers where they can find the info you are citing. But if that info is unpublished, then that's life.
Why should this not apply everywhere else?
E.g., Stripe citing Paypal for their idea, Google citing Altavista, and Facebook citing Myspace.
As a side note, why do you believe that it doesn't apply outside those industries? I'm sure if you asked Brin and Page, they would say they were inspired by (the limitations of) previous search engines. Ditto for Stripe and Paypal and Facebook and Myspace.
Academic ideas are (hopefully) much more specialized than these broad ideas. In business they are more comparable to patents and for patents prior art is of course extremely important (enough so to legally invalidate your legal claim to your idea!). In contrast to patents, however, there aren't strong legal mechanisms to enforce priority of ideas and that's why being honest and open about it (community policing of these ideals) are especially important. If academics were to entirely stop caring about citing their sources, academic research would probably totally cease to function.
>A large number of seminal works have never been published. The greatest mathematics paper of our lifetimes remains unpublished.
Is this a reference to a specific paper?
I should note, since I am not and do not expect to be the level of mathematician that Perelman is, I have not actually read his proof. So I defer to other superior mathematicians for this assessment and come by it as hearsay. :)
There is actually two aspects about "published". One is archival, so people can expect to access the work decades later (if they have to pay for that is another discussion). The second aspect is peer-review aka quality control.
Personally, I once submitted a paper to a workshop. After submission, peer-review, and acceptance the workshop committee decided that they will not publish proceedings. I could have submitted the paper elsewhere, which I find weird. Instead I published it as a techreport. However, it is now unusable for proposals, because a techreport is "not published" even if it is properly archived and went through peer review.
IMO, the definition of 'published' is a huge issue. I have always read published as archived peer reviewed research. Your paper is both, but remains in an not published state which I think is wrong and hinders future research.
[1]: https://arxiv.org/abs/1703.04933 [2]: https://arxiv.org/abs/1701.07875 [3]: https://arxiv.org/abs/1706.01350 [4]: https://arxiv.org/abs/1512.04860
That said, everything I saw in the papers you linked was linear algebra, calculus or probability theory plus the usual smattering of background notation and set theory.
Once you have a solid background in those areas, it is likely more productive to look up the specific concepts mentioned in a paper (such as the Kullback-Leibler divergence or the Bellman equation), because by then you are probably too deep in the woods to find one resource that adequately covers all those different directions.
Books are probably a less efficient method of learning the mathematics if you have targeted subjects you want to learn about. They're typically suited to introductions and breadth-wise coverage of fields, but once you get higher up, "linear algebra" (for example) can get fuzzy with things like abstract algebra. That means you'll end up with several tome-like books to work through which can be productive, but it'll take a while and you'll need to map the material to the applications you're interested in on your own. It's more efficient to develop a good baseline of understanding about a broad subject area, learn the foundational theorems, then move on to the specific areas you need to learn. This is typically doable if you've developed the requisite mathematical maturity overall and if you have learned the "essentials."
Practically speaking: maybe pick up foundation texts like Strang's (linear algebra), Spivak's (calculus) and Ross' (probability theory). You're going to want a solid foundation in analysis before moving on to higher order probability theory, so drill down on that after you do a refresher on the calculus. From there you should attempt to read each paper (even if you struggle a lot), take notes on what confuses you or doesn't make sense, read the prior art on those topics and then come back to it.
I don't particularly read machine learning papers often, but I read mathematical cryptographic ones very often (at least once per day I find myself in a new one). It's not typical that I read a research paper introducing a novel primitive or construction where I follow the math immediately on a single pass, and I often come across things I need to read about first. From a thirty thousand foot view the math for both of these subjects is broadly similar in rough topical surface area, so I think this methodology for academic reading is fairly applicable to most subjects that involve a lot of mathematics understanding.
Basically: don't approach learning the heavy math with a monolithic, brute-force approach as if you were in university. That's a slog and it's demotivating. Learn the minimum foundation for each area you need, then proceed to more advanced topics as you need them.
EDIT: I think the main topic missing from my background is this so-called "analysis". I never formally studied it. Is there a more efficient way to study analysis than spivak's, for someone who has a decent background otherwise?
Analysis is basically "really rigorous calculus". Basic analysis courses are also usually where you learn to do proofs.
(To some reasonable generality "calculus" stands for "rules of manipulation", while analysis is the mathematical theory of calculus. So I can teach you stochastic calculus in a couple of two-hour sessions but understanding what the hell is going on (stochastic analysis) requires measure theory, some functional analysis and much courage)
Where do you find new ones? I'd like to get into this.
The linear algebra is obviously key, and I wish we'd done more of the in my advanced high school classes instead of elementary analysis.
There's a similar issue in journal publishing: counting "published works" without regard for where leads to journals that will publish literally anything for cash.
Is this really a thing?
I understood that today you even city web pages (with a time stamp) and even your dog.
I wish there was some way to generate all relevant bib information as soon as paper gets accepted which then can be added on arxiv immediately. This would allow folks to distinguish between peer reviewed papers vs those which are submitted only for flag planting.
I adore arXiv but still believe it's a preprint. In my field (cognitive science) it would be great if we had more methods to sidestep Elsevier and the other commercial publishers and have an open stack with rigorous peer review (PLoS being the main way currently).
There is a fair amount of evidence that this isn't true. In general, most statistics in scientific research aren't done by statisticians, and there are whole classes of methodological errors that are regularly not caught because the "peers" have the same lack of statistical education as the people whose papers they are reviewing.
I see that as all true, but irrelevant in this context. If you source material from a pre-print on arXiv, then you should cite it. Seems totally obvious to me. Of course you would prefer the final, published paper if it's available. But that wasn't the question at hand.
And even with all that said... I would argue that in some fields, (cs / ml / etc.) we're getting close to a point where arXiv itself is become almost a parallel publishing mechanism where people cite/publish completely within the arXiv realm, with less regard for "traditional" journals and what-not in general. Especially when you factor in papers from researchers who come from industry, as opposed to academia, and care less about some of the normal trappings of academic publishing.
I adore arXiv but still believe it's a preprint.
Of course it's a pre-print. I didn't contend otherwise. I'm just saying that, from my perspective, it's obvious that you should cite a pre-print if it's relevant.
I will allow though, that norms probably vary from field to field, and as a non-academic, my take is likely different from, say, somebody who is deeply immersed in academia, pursuing tenure, etc.
- novelty
- technical correctness
- clarity
- good experimental evaluation
Novelty is the most important: if without the citation your paper looks like the original idea, this is plagiarism.
If the paper has big shortcomings that your paper adresses, it is fair to give yourself the credit you deserve of course, but it doesn't harm to cite the other paper, in fact it gives a way to give some sort of peer review: In [1], Foo and al. attempted to explore <subject> but the experiments were inconclusive/the technique sucked compare to state of the art/they didn't explain how they did it... In this paper we did this and that and it gives us awesome results (Said in a nicer way)
I think you should cite any idea you pick from a paper, full stop. That's the whole point of citing: To assign credit where credit is due.
I've been told a rule of thumb is that if a related unrefereed arxiv paper has been cited six or more times, with the justification being that this means it is somewhat well known once it has some citations.
It definitely does not matter how many citations the other paper has. The point isn't to avoid getting caught, it is to inform the reader. Your citation is more useful the less well known the cited paper is.
I can see how this would be maddening, particularly if you started before the flag-planting paper was even written.
> Yes, of course. Any time that our work follows, copies, or borrows ideas from other people, and when we can reasonably be expected to be aware of this, we ought to cite the related work.
> We should not have to cite nonsense. Many reviewers are abusing the system and asking for ridiculous comparison to recently-posted preprint papers. Bald-faced flag-planting should not be rewarded. And we should not be faulted by reviewers for failing to compare against 2-week old algorithms that may or may not work.
So what position is the author advocating? Citing or not citing?
edited
Something being on arXiv or in a blog post or … doesn't excuse not citing it if it influenced your work. It's important to document where your ideas and data come from, both to give credit to the author and to allow others to evaluate what you base your claims on for themselves.
The second quote is about stuff that didn't influence your work. While you are expected to keep up with and document related developments, forcing authors to constantly update references to new, not yet properly evaluated work just because it makes some related claim doesn't make sense.
It doesn't say that though, it just says that you shouldn't have to cite nonsense on arXiv. This implies that you can read a paper, implement something similar, and afterwards decide that the paper was nonsense, didn't influence your work, and shouldn't be cited.
The context seems pretty clear to me. Your example clearly is covered by the first case: if it influenced your work, you cite it, even if it is "nonsense". You can't just "decide" something didn't have influence if it had. (These rules do not prevent cheating, they are guides for people acting ethically)
The second rule is to prevent the opposite case: You shouldn't be forced to create the impression your work is based on or just a mere repeat of someone else's "who had the idea first" when they have no good claim to that, or inferior to something that hasn't been shown to be actually better.
The question here is whether you should cite something which you did read, and does relate to your work, even if it's a shitty flag-planting paper.
If you think it's shitty, then you can cite and dismiss it in a sentence. You can dismiss 30 papers in a single sentence if you like. There's no requirement to wax lyrical for 3 paragraphs about a paper just because it was first. But it strikes me as dishonest to advocate that sole researchers become arbiters of a paper's merit, citing or not citing it at their personal discretion.
Also, if their paper was published first then that's the only claim necessary to demonstrate that they were first out with the idea.
> Also, if their paper was published first then that's the only claim necessary to demonstrate that they were first out with the idea.
To quote myself: impression your work is based on or just a mere repeat of someone else's, not just being first. Ideally, everyone looking at you referencing it would take note that it was published months after you started work and your work was independent (or even earlier), but that easily gets lost.
Should you be encouraged to throw out every idea and snippet to arXiv just so you can claim "FIRST!" in case it turns out to be useful/true, over "competing" works that spent more effort on quality and verification and are now in peer-review forced to reference you as the pioneer (even if you maybe had the idea months later, but rushed it out and got lucky with it holding up)? That's what the "flagplanting" is about.
On the one hand, if you borrow an idea from another person or source, then you should cite it. Just because it's only the arXiv doesn't give you a pass not to.
On the other hand, "flag-planting" articles are not something that you borrow from so you don't have to cite them - and you shouldn't as the practice should not be rewarded.
As `CogitoCogito points out, it's routine to cite completely unpublished material such as private communications.
Conversely, if a flag-planting article somehow makes it into a very prestigious journal, then you can still ignore it.
So really, the publishing status is only an initial filter, and the potential source should always be judged on its merits.
If it doesn't relate to your work, or you haven't read it, then obviously it needn't be cited.
To take an example, read the article linked to in the OP's. The author describes how terrible two papers are (implying that they're not worth citing) only for one paper's author, and other researchers, to come on and tell him why he's wrong about his interpretation and understanding. This leads to him retracting his claim that it wasn't worthy of merit.
So immediately you have an example of a sole researcher deeming himself to be the only judge of merit necessary, only to be wrong. His judgment has a 50% failure rate already, and that's with him cherry-picking 'bad' papers.
This position can't be advocated.
If you have chosen to publish on arXiv, then you have chosen to step out of the peer-review route. The author of the linked article did not deem himself to be the only judge of merit necessary, his position as sole reviewer came about through the decision of the papers' authors to publish on arXiv, and it seems the article's author would have preferred it if the papers had been well-reviewed before publication. You are not advocating for arXiv papers to be immunized from evaluation, are you?
It seems that we are rediscovering why the peer-review process, with all its flaws, was created in the first place. Complex problems rarely have simple solutions.
I'm in agreement that ideas are a dime a dozen. Sure, it's necessary to cite the first known mention of an idea, and there are situations where the first mention of an idea is important such as in the patent system.
"Planting" happens in my world all the time. Unfortunately, managers give a lot more importance to "ideas" than they are really worth, because they over-value their own interventions in general. Somebody will blurt out an idea in a meeting, wait until someone else has developed it, and then rush in to take credit. If a manager does this, it's a blow to morale. I have my own rule of thumb, which is "show your work" from math class. Just writing down the answer doesn't get you full credit.
A couple of historical examples: The ancient Greeks are credited with the atomic theory, but they had no concept of even turning it into a serious hypothesis. Lots of ideas are anticipated in science fiction, but do those authors really deserve credit?
I see what you mean, but this is mostly office politics, which is not very related to the academic citation process. If an author makes a conjecture that stimulates further work (even if it's just on the ArXiv), he deserves to be cited.
> Lots of ideas are anticipated in science fiction, but do those authors really deserve credit?
Well, I guess it depends on how much the idea is fleshed out. Lucian was probably the first author to conceive space travel, but it was just "people land on the moon" (still pretty far out for his times, though). Meanwhile, Asimov's three laws of robotics depict a reasonable control scheme for e.g. an autonomous vehicle, so if they are somehow implemented their author should be credited.
In my view, using citations as a metric, and abusing that metric, is just an academic version of office politics, writ large.
Polywater is the perfect example of why you should citate. Just because someone publishes something does not mean its correct, scientific, or proven.
https://en.wikipedia.org/wiki/Polywater
http://science.sciencemag.org/content/167/3926/1715?sid=8b4e...
Science takes time.
Write a good paper and cite, cite, cite! If your citations are ever demonstrated to be incorrect or fraudulent other researchers can continue work to disprove the citations and work on correcting the errors.
Given the amount of information available, it could often be the case of independent research into something which is known and available for some time. If one only learns about similar - and possibly greater - results after making one's own, and wants to talk about the work done - should one cite other, possibly earlier, works?
If you really believe in science then ONLY non-paywall papers should be cited.
http://approximatelycorrect.com/2017/08/01/do-i-have-to-cite...
Anyway, I didn't mean to argue. Have a nice evening :)
When there are two articles with the tag, they will both be shown on the 'publishing' page, in their entirety, newest first. See [1] for an example.
Redirecting would be akin to Google redirecting you to the current top result (like I'm feeling lucky, but for all searches)
The project is to make an FPGA implementation of a technique presented in an arXiv paper. The paper had some big gaps, so I had to spend a lot of time researching the technique, and I had very little time to spend on the actual implementation.
In CS that almost certainly isn't true. I'm most familiar with the NLP field, but there, if you have some kind of embedding of your words/tokens/sentences/something you cite https://arxiv.org/pdf/1301.3781.pdf (Word2Vec, Mikolov).
That paper says there is a follow up paper published at NIPS2013, but I don't think I've ever seen that published.
The field just moves too fast to wait for conferences anymore.