Researchers discover a new form of scientific fraud: 'sneaked references'
phys.org
phys.org
> When registering a new publication and its references at Crossref, a publisher may sneak extra undue references in the metadata sent in addition to the ones originally present. Then, digital libraries (e.g., SpringerLink) and bibliometric platforms (e.g., Dimensions) harvest these metadata, undue citations included. These sneaked references are processed and counted even if they are not present in the original publication.
The three journals in this particular case are all published by Technoscience Academy, an OA publisher operating out of India (not one of the well-known ones). I would think twice as an author before I submitted to any journal from this publisher, lest my paper is abused for manipulations like this (although I'm not sure if it has any journals worth submitting to anyway).
NB (because I got confused first): This is not really about Hindawi. Hindawi published the (trash) article that these fake citations were pumping up, but the pumping-up happened using Technoscience Academy journals.
In recent years payments based directly on the number of citations a paper receives have become more popular, but are still much less common than those based on the journal’s impact factor.
https://opportunities-insight.britishcouncil.org/insights-bl...
If you make the analogy between the www: H-index is pagerank, citations are back links, and authors (researchers) are the domain names. Gaming h-index is akin to SEO hacking for academic authors.
“The petitioner provides evidence demonstrating that the total rate of citations to the beneficiary’s body of published work is high relative to others in the field, or the beneficiary has a high h-index[30] for the field. Depending on the field and the comparative data the petitioner provides, such evidence may indicate a beneficiary’s high overall standing for the purpose of demonstrating that the beneficiary is among the small percentage at the top of the field.[31]”
So Google Scholar uses the text which is good. Then obvious solution is to go and look up where it has been cited, which would be easy to do with google scholar.
I don't know why anyone would jeopardize their career like this.
Publish or perish.
Is this different in other fields? Or in sketchy journals?
> For example, a single researcher who was associated with Technoscience Academy benefited from more than 3,000 additional illegitimate citations. Some journals from the same publisher benefited from a couple hundred additional sneaked citations.
Perhaps this publisher or others also offer this as some kind of backroom deal / service.
It's possible that some of the inconsistency between metadata and text could just be due to incompetence - it's harder to find a profit motive for dropping legitimate citations. Why wouldn't this sort of metadata auto-generated from the text (aside from enabling fraud, of course)?
Competitiveness for citation points, especially with someone in or adjacent to your niche?
Also, the non-profit: pettiness.
In the first example shown in the linked pre-print [1] there's a paper with 62 downloads that's been cited 107 times within two months. The pre-print looks deeper into a paper with 7 "real" references whose metadata has an extra 40 references not found in the PDF. This leaves us with three options:
* the author of a paper with 62 downloads (not an amazing number) was convinced into joining a citation ring along with 40 other authors,
* the publisher has been sneaking references onto unsuspecting papers, or
* the publisher has a vulnerability on their metadata system that's being actively exploited by the two scholars identified in the pre-print.
Whatever the case, I'm glad the solution is as simple as "you should parse the references yourself". I do however wonder: is someone checking whether all of the references are actually referenced within the paper?How would that work for paperback references? That would be a nightmare. If an author cites 20 different sources, verifier needs to checkout 20 sources at the library (if they are even available)
They should be able to find citation "rings" then, whole groups which regularly do this, probably associated with specific institutions or journals.
The linked study did part of this: https://asistdl.onlinelibrary.wiley.com/doi/10.1002/asi.2489...
> An analysis of the 10 sneaked references in Dimensions reveals that they benefit mainly two authors (Initials JNR & BK)
Now, it would be interesting to see if JNR and BK's publications used this trick and in turn benefitted, some other group.
This is a problem with the journal review and editors. Also, typesetting tools that create the final version can and should be setup to protect things like these. I know folks may want to go hunt for sexy genai tooling to solve this - but I think the solution is much simpler.
> This is a problem with the journal review and editors. Also, typesetting tools that create the final version can and should be setup to protect things like these. I know folks may want to go hunt for sexy genai tooling to solve this - but I think the solution is much simpler.
The issue in the article isn' a paper being listed in the references but not actually cited elsewhere of the paper; it's not something within the actual paper at all. It's metadata created by the publisher.
So it presumably doesn't have anything to do with what latex allows or doesn't allow.
This is probably the publisher's doing rather than the author of the paper.
Wouldn't surprise me in the slightest to find they demanded authors create/submit the metadata with the references in it, and that it never gets shown to the unpaid peer reviewers, and is never checked by anybody.
Instead, the published article should really be a view of the structured data, metadata, and text (i.e. the true content) that makes up the article. Formatting and such can be a pain, but using this approach would mean the published article is a view of the truth rather than the metadata being created as something of an afterthought.
These additional references were only in the
metadata, distorting citation counts and
giving certain authors an unfair advantage.
Papers with metadata that doesn't match the contents of the paper. The article notes that Google Scholar is unaffected, as it extracts citations from the paper itself by parsing the text of the printed bibliography.The problem would be if this turns into a negative index, it can have equally bad repercussions. So attribution to malice and intent is important because there will be people adversely affected.
If this is a publisher/SEO fuckup, that needs to be seen distinct from "fraud"
https://cadenaser.com/castillayleon/2024/03/15/el-candidato-...
Sometimes there are just five people working on a particular theme. Avoiding entirely to cite your previous work because some people could frown is stupid. Hides a fifth of the extant knowledge to the readers for no real benefit. As researchers now are required to produce constantly, they only can release small increases in their work. Each article isolated will not made any difference, but are part of a slow chain. Results don't come necessarily in a constant predictable stream. Citing the previous chapter in those science series is reasonable or even necessary to understand the current article.
Citing only your work and not other's work would be the problem.
And sometimes this things just happen. I could talk about "my curriculum" on the official web of my university. A list of several articles with my name. I never wrote, toke part on the design on the web or supported it in any way.
And is totally fake. I didn't wrote a single one of the articles cited on it.
After scratching my head a little I see that they commit a obvious mistake in the database query (that is not my problem anymore to fix).