Dozens of scientific journals have vanished from the internet
sciencemag.org
sciencemag.org
https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3...
and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ, ISSN, DOI registrars (Crossref, Datacite, others) are crucial for this. In the broader ecosystem, we hope this can complement existing efforts that partner with large publishers (like LOCKSS, Portico, JSTOR) and institutional repositories. A natural niche for us is web-native (HTML) content, which we have crawled a lot of but are just getting started to index. For example, publications like d-lib, first monday, and distill.pub.
If folks want to help, it would be great to have a "youtube-dl for open access papers". There is a lot of content on large platforms and publishers which have anti-crawling measures (even for gold OA and hybrid content!), as well as a long tail of small publishers that don't use simple/common mechanisms like OAI-PMH and the `citation_pdf_url` HTML meta tag to identify fulltext content. The OAI-PMH ecosystem sadly is not very complete or helpful for the use case of mirroring.
Zotero has an existing set of "translators".
And somewhat related to this request: I scratched out some notes last year about how to get more out of "zero-obligation communities" (like the pool of prospective contributors in open source) <https://www.colbyrussell.com/2019/02/15/what-happened-in-jan.... The long and short of it is that instead of saying something like "if folks want to help, it would be great[...]", you should provide a place for people to sign up, take them at their word that they're willing to help, and then lay out a concrete set of tasks/deliverables. People get weird about trying to avoid being seen as not gentle enough with volunteers, but the end result is a lot of unharnessed human potential. You've got a pool of mechanical turks at your disposal. Take a break from polishing the arrangement of instructions you give to the computer and focus some energy on writing the "programs" that you want to be executed by meatbags.
Most of orgs/journals/conferences in this group just don't have the resources to be able do this, nor maintain something like this.
It was funny last year when a rep from google was doing a teleconference talk from Mountain View in Jakarta, covering things like this in front of a room full of like the top 200 journal managers in Indonesia (there are like tens of thousands of journals) and their eyes were glossing over when the rep was trying to address fixing some of the issues they encounter when crawling than hamper just indexing.
Might as well be coming from a different planet…
The trickier cases are when folks write their own platforms, or even write their own raw HTML with no templating, in which case adding tags to all landing pages or supporting an API would be a relatively large amount of work.
Yeah, and there are a lot of issues with OJS and the meta tagging and google being able to crawl a lot of these (not to mention site uptime where lots of sites go down for long stretches and google just assumes the site was taking offline permanently if kept down for a while). Esp when the meta tag locales don't map to the actual language used in the papers themselves (i.e. journal admins enter meta data tagged as EN but use indonesian for the text and the text of the paper) or like non iso locale codes in metadata fields, etc.
Right now, when journals/orgs/conferences want help getting indexed in google and have oai endpoints, we pass all their content through locale detect stuff and don't trust their meta data by default.
I think Unpaywall is already trying to do this? Or at least, for every DOI they index, they try to include a link to the direct article if known.
Having an URL to something that is supposed to serve a PDF doesn't mean that you actually get it. Publishers like Elsevier or Wiley require a cookie/JavaScript dance and/or place heavy rate limits; whatever PDF hosted by them is never really accessible, even if it's ostensibly open access.
The "youtube-dl for papers" would do what youtube-dl does, i.e. take the URL and execute whatever JavaScript or other trickery the target page/website requires before it serves the content.
https://en.wikipedia.org/wiki/Hyper_Articles_en_Ligne
"Hyper Articles en Ligne, generally shortened to HAL, is an open archive where authors can deposit scholarly documents from all academic fields."
I work at a university and I know library people and management check carefully that every paper we produce is deposited in HAL.
Also:
https://fr.wikipedia.org/wiki/Hyper_articles_en_ligne
"Depuis le 25 septembre 2018, les dépôts de logiciels sur HAL sont connectés à Software Heritage"
https://en.wikipedia.org/wiki/Software_Heritage
For recruiting some french institutions like CNRS will only consider papers deposited in HAL when doing the evaluation.
However, most of the archive is not publicly accessible due to copyright, privacy etc. You can request access to specific content both as a private person and as a researcher.
https://pubmed.ncbi.nlm.nih.gov
It works pretty well. Papers are submitted. The get a unique id. Some are stored and accessible. (Some US funding sources require public accessible papers.).
NCBI has a ton of information. Its a pretty awesome resource. https://www.ncbi.nlm.nih.gov
They even index the paper submitted with a controlled vocabulary of terms (within a month or two) https://www.ncbi.nlm.nih.gov/mesh/
However, almost 12 years of tweets is still very cool. A LOT of startups probably had quite a bit of material in those early days of social media.
But today, the institution announced it will no longer archive every one of our status updates, opinion threads, and "big if true"s. As of Jan. 1, the library will only acquire tweets "on a very selective basis."
The library says it began archiving tweets "for the same reason it collects other materials — to acquire and preserve a record of knowledge and creativity for Congress and the American people." The archive stretches back to Twitter's beginning, in 2006.
The institution says it will continue to preserve its collection of tweets from the platform's first 12 years, but indicates that it has yet to figure out exactly how to make the archive public.
https://www.npr.org/sections/thetwo-way/2017/12/26/573609499...
But all this misses the point a little -- it is not just the articles that should be preserved, but the journal itself, as a collection of articles with metadata (including the fact that it was collected in the journal), records of editorial boards, editorial articles, etc.
However, arXiv is (or aims to be) a supra-national organization. Do you think it is preferable to have an international standard repository for this knowledge like arXiv, a network of federated, national systems such as HAL, or both?
ArXiv has historically sometimes found itself hard-up for funding, so I think that a valid argument could be made for both approaches.
[0]: https://arxiv.org/
Which is strange, as you paid them to publish your article, that was likely paid for by state funding -- aka tax money. Ah, the academic racket is so beautiful.
"Transfert automatique des documents vers une archive ouverte internationale telle qu'ArXiv ou Pubmed Central "
It’s totally counter to its original purpose, and was implemented for researchers to be able to fulfill the internal CNRS mandate to list their paper on HAL without needing to actually host the paper there.
For instance, there are about 140k works under a Creative Commons license (mostly free ones, some unfree) from French authors, most of which aren't preserved on any open archive. https://www.lens.org/lens/search/scholar/list?p=0&n=10&s=ref...
While I like HAL in general as a consumer, that policy is terrible if true for people like me who started their academic career abroad. One more reason not to go back, I guess.
There is actually a tool that is supported by HAL that makes it easier to quickly add all your publications by just giving your name: https://dissem.in/
For instance, it automatically fills in all the metadata for your publications.
Work funded by the NIH needs to end up in PubMed Central within a year of publication. The NSF and DoE have similar policies. Unclassified DoD-funded work also needs to end up in the Defense Technical Information Center. All of this flows from a 2013 memo by John Holdren/OST entitled "Increasing Access to Results of Federally Funded Science", and, as far as I know, it hasn't been overturned. While this doesn't formally cover everything, it comes pretty close and many journals now handle this automatically.
The memo is here: https://obamawhitehouse.archives.gov/blog/2016/02/22/increas...
Canada has a similar policy for Tri-Council funded research (here: http://www.science.gc.ca/eic/site/063.nsf/eng/h_F6765465.htm...) but does not require specific repositories.
list of the 176 vanished is here: https://github.com/njahn82/vanished_journals/blob/master/Dis...
As someone who hates to see this stuff disappear, there is still a cynical person inside me that knows there are a glut of journals that are often used to bump publishing count for professors trying to get tenure.
That cynic inside of me also realizes that a subset of the journal business is a bit of a scam anyway since frequently authors have to pay the journals to include their paper and then the journals charge an exorbitant rate to get access. Have you tried to buy a journal article lately ($35 for one paper!) -- yeah, neither have I.
There are multiple works to show that the rise of search engine like Google Scholar have meant that researchers are increasingly citing the same papers, because their searches are all returning the same thing.
Meanwhile, there are some "sleeper" papers that are super relevant to a lot of works, offer great insight, but by virtue of their low search ranking, never get cited.
That's not to say there isn't a fair amount of unremarkable research. It's just that it doesn't always correlate with citation count.
Curiously I have another co-author paper by another student that was published around the same time (same topic different application) and as of today it has over 60 citations count. It even cited by a Stanford University researcher but only with classification accuracy of below 90%!
I think this is a classic case of a sleeper paper where you have excellent results but then nobody want to cite you because it makes their paper look bad or negate their (fake) claim of novelty. FYI, most of the papers in the field use the same venerable open online database that make it easy to compare your results against others.
1. https://en.wikipedia.org/wiki/Predatory_publishing
2. http://web.archive.org/web/20170128194124/http://www.chronic...
Perhaps it's arguable whether that is to be called a scam, but certainly it is a rip-off.
1. cheap it is, considering the size of the institutions that produce and benefit from them.
2. absurdly broken the private publishing industry has become.
Hum... What?
Universities at worst don't care. Most really want they work circulating and will do a lot of things to get it (many useless things that miss the point, but well, that's how people are).
Universities could push it harder. But they are surely pushing on the correct direction.
I wonder how aaronsw would feel about this statement.
"Very early in this post-arrest period, MIT decided to “remain neutral,” as between the government and Aaron Swartz, in the investigation and eventual prosecution. Initially this meant simply that MIT would not take a public position on the prosecution.Throughout the following (almost) two years, MIT’s decisions were mostly guided by this posture of neutrality."
"With regard to substance, MIT would make no statements, whether in support or in opposition, about the government’s decision to prosecute Aaron Swartz, the government’s decisions about charges in an indictment, or any possible plea bargain stances of the prosecution or the defense. [15]"
"[15]: This position of neutrality would not have necessarily extended to the sentencing phase of the prosecution, where MIT might have been prepared to advocate on behalf of Aaron Swartz had he been convicted."
My point was admittedly an aspirational one that requires major changes to the current IP ownership model of (often publicly funded) scholarly work. For example: see how publishers have used the legal system to throw the book at projects like sci-hub or individuals like aaronsw.
I do disagree that the main issue is universities not being keen on having their work circulated openly.
My take is this:
1. Publishers used to have a reasonable value prop, but are now basically robbing the public blind.
2. University admin roles, even at public universities, are more and more occupied by metric-obsessed, run-it-like-a-business types leading to:
3. Academics are stuck on the publish-or-perish hamster wheel that is fueled by a vague amorphous notion of journal prestige, set up and profited by said publishers.
I can't see how we can fundamentally change this situation without either of:
- top-down structural change in academic admin system allowing people to pursue academic careers without h-index and citation obsessions. Major tenure reform?
- bottom-up mobilization of academics or public interest groups, forcing publishers' hands into a more sane (and less profitable) business model. They will fight it tooth and nail of course, hence the need for mass coordinated action (boycotts, public shaming, etc.)
Both these avenues need playing politics on one level or another.
EDIT: formatting and clearer wording.
This second points is important for having a coherent discussion of the literature and avoiding fraud to some degree, and to their credit, the major journals do provide this (but not the first point). In this sense, I think Pubmed Central is an example of a good way towards a public utility model:
Take arXiv and expand its scope and upload everything there. Boom, problem effectively solved using mostly existing tools.
https://www.ncbi.nlm.nih.gov/pmc/about/intro/
Most NIH funded work is required to be deposited there, but at present, some journals place an embargo on new papers, giving them a year before they are available on PMC.
AFAIK the problem that the OP points out would not have been fixed by Pubmed. It would keep the abstracts, but the full contents would become dead links because the publishers _own_ the full article at the end of the day (which is absurd IMO).
And for archival: https://www.lockss.org/
What's the digital, and post-national solution?
It is somewhat trivial to devise an API to be integrated in the publications pipelines to automatically and transparently submit new and modified articles to a central repository.
It is, as far as I can tell, mandatory in theory but not in practice. https://www.copyright.gov/help/faq/mandatory_deposit.html#:~....
I don’t think the solution is to move back towards the old model. There are already lots of initiatives towards creating online archives of academic work that may be piggybacked on. In mathematics, perhaps the easiest way to set up an open access journal is as an arxiv overlay journal where at the most basic level each issue of the journal is a list of links to specific versions of papers on the arxiv. This would be likely to be archived sufficiently well.
For a traditional journal that shuts down to be archived, lots of things need to happen:
1. Some library needs to pay some exorbitant fee to get physical or (permanent not saas-based) digital copies of the journal
2. That library needs to keep hold of that copy for the 100 years or so until copyright expires
3. That library then needs to take the initiative to make its copies available
This seems like a harder process than finding some public domain digital copy. And for a lot of journals, the only reason the library gets copies is due to the bundling systems which universities hate.
I’m curious to know more about these journals which did vanish, and what sort of quality they are. If a predatory journal offers open access and later disappears, would they be counted?
I think they legit view preservation of the scientific record as within their provenance to cover.
LOCKSS provides a free option for publishers to join, but only accepts a limited number of OA publishers (https://www.lockss.org/use-lockss/publishers). A couple years ago the PKP launched their preservation service, which we're really excited about as it also offers free preservation (for OJS journals) and would help esp. those smaller journals that otherwise couldn't afford to enroll into preservation schemes.
I don't really agree that traditional publishers lose all articles by default by this metric. I think there is some value in a reliable record of the journal itself, as more than the sum of its parts.
Still, my preferred solution for all of this is like JMLR -- very low-cost and open access, and it has reliability by virtue of association with a top university, and prestige by virtue of its editorial board (which becomes self-fulfilling).
Having archives is so important in the legal field and plenty of areas of research.
Very good point. I will donate to them too. I've been donating to Archive.org for a long time but I use Archive.is more often these days so they deserve some love too.
I have a list of places where I make donations to on my profile page here.
Well, this exact definition could have some false positives, like a journal that publishes every third volume as complete open access and keeps the others behind a paywall. But I'm sure they were a bit more careful than it says here.
In some cases, we found websites (other than the original journal website) that now host some individual issues, but not all of the content.
For this reason, customers (mostly libraries and universities) around year 2000 have started demanding that the closed-access publishers have preservation mechanisms and provide dark archives to ensure access after subscriptions expire or are otherwise breached.
This has resulted in initiatives like https://www.lockss.org/ and https://clockss.org/ .
> Since CLOCKSS launched in 2008, 53 journals comprising 13,000 articles have been triggered
That's 13k articles (mostly from closed-access publishers) saved from oblivion, but there's many more. The English Wikipedia alone knows thousands of articles where even the DOI is broken (usually they're from publishers like Elsevier, Wiley, T&F, LWW, OUP). https://en.wikipedia.org/wiki/Category:Pages_with_DOIs_inact...
On average, fully open access journals present a much lower risk of vanishing because they're easier to archive, especially if they use a [free license](https://en.wikipedia.org/wiki/Free_license) like CC BY or CC BY-SA. However, publishers still need some nudging towards archival.
DOAJ requires a digital preservation plan in place (at least with CLOCKSS, LOCKSS, PKP PN, PMC, Portico or a national library) for a journal to be granted the [DOAJ seal](https://doaj.org/publishers#seal). At the moment there are [about 1400 journals with the DOAJ seal](https://doaj.org/search?source=%7B%22query%22%3A%7B%22filter...).
Authors can publish in a journal with the DOAJ seal and be sure that their work is safe. On the other hand, they have no recourse against a commercial and closed-access publisher, which can be forced to comply with archival only by contracts with its paying subscribers.
Anyone disagree?
1. The importance of a piece of scholarly work need not be immediately apparent to its contemporaries.
2. Being cited is a poor proxy for importance.
3. Work that is cited is rarely paraphrased or duplicated in a meaningful way.
4. Paraphrased citations are a poor proxy for canonical source; papers are often cited incompletely or sometimes outright inaccurately.
5. These are the fruits of people's labour. They spent days and months producing them. To lose them, especially when digital copies are so cheap, is an unnecessary disregard for said labour.
What do you mean by "duplicated elsewhere"? Sitting on some scientist's hard drive doesn't count, it has to be discoverable. That's what the issue is: when journals die, how do we ensure that the papers are saved somewhere easily searchable?
Meanwhile people are working hard to preserve every Commodore 64 game.