I'm a programmer – how can I help SciHub?
reddit.com
reddit.com
https://github.com/subdavis/libgen-seedtools
The sci-hub archive is partly supported by libgen.rs. To ensure that their content remains accessible, they have thousands of very large torrents, many of which are not well seeded. If you have a few TB of disk space and bandwidth to spare, it's a good way to help out.
How does that _actually_ work, though? Can anyone download the final, published PDF from "publicly accessible repositories"?
(Full disclosure: ex scientist with published papers, still no idea how I can legally share _my_ work with anyone who might be interested...)
http://lesscrime.info/post/how-to-stop-hiding-your-research/
You might be right about the far future, but there's still a lot of human flourishing that fails to happen every day until then. Or you could be wrong, and I'd hate to /start/ having this conversation once we realize that.
Yes, but the perfect can be the enemy of the good.
Personally, I think part of the solution would be to have grants that are in some way publicly funded (taxes) to have a portion set aside to pay publishing costs, and require publishing in some way. This would both make open access with well-polished articles more accessible, it would also help solve the issue of negative results rarely being published.
Not perfect, not a.silver bullet, but at least an incremental improvement.
Claiming $10k for a bunch of unnecessary work is outrageous.
The cost of prestigious journals are the curation and high standard of peer review. Except they don't actually pay their staff for those things...
But you are right: it's still a mostly unnecessary cost these days. And regardless if the other costs, $150 to $700 per year for a journal subscription is rarely justified by the publishing costs the turn articles into more polished and high quality pieces.
The financial model for disseminating scientific knowledge is broken, but I did want to acknowledge that there is good and useful work done in-between an author writing up their research and ultimately making it available online.
I'm assuming journals were happy to switch subscriber to digital-only subscriptions. All the revenue without the circulation headaches.
In the very early 00's a small publisher might have a few dozen journals and a dozen books a year, and reasonable but not huge profit margins. At the same time a couple of things were happening:
1) Elsevier was already snatching up small publishers and offering their own digital access.
2) For a small publisher going digital was much more difficult. It wasn't in their existing tech skill set and there weren't much in the way of turn-key solutions to fill their specific needs
3) Royalty agreements weren't built for all-access subscriptions. They were based on purchases of a full issue. They couldn't be shoe-horned into all-access subscriptions so the publisher would have to go back and renegotiate contracts for digital distribution in a process that was extremely time consuming... getting agreements for 30 years of your back catalog wasn't easy.
4) For a small publisher, digitizing content was very expensive. A simple destructive scan where the content was unbound and image scanned to a PDF wasn $0.50 to $1 per page, but you need things to be searchable too. OCR wasn't good enough and OCR w/human correction was very expensive. Common practice was to go for the cheaper image-only OCR and tack on ToC, Titles, abstracts and authors as indexed metadata, which also cost more money.
Elsevier was big enough that most of the above were somewhat non-issues for them. They could throw their bulk around and dictate terms on royalties deploying an army of people to work them all out journal by journal. They had plenty of money to build a digital platform and digitize their back catalog year by year at a fairly rapid pace.
Library budgets were being cut back, so it looked like a real nice option to subscribe to Elsevier and get instant access to tons of stuff. It became a lot harder for smaller publishers to get their material into libraries. This created a downward spiral, Elsevier picked off small publishers one by one, making their product even more appealing.
In the end, we're left with what we have now: A few massive publishers that looked like reasonable options 15 and 20 years ago at those prices, but then the prices rapidly escalated, and after 7 or 10 years of digital subscriptions you almost couldn't go back because you'd have to buy that many years of back catalog.
Uhm, I don't know if you've published a paper since the 90s but none of those are costs that modern journals incur. Or if they do, nobody asked them to.
The main thing journals do is peer review and that is all done for free by other academics. Authors do basically all of the typesetting, and although journals still insist on printing issues there's really no need for them to do so.
The only really important thing that journals do is finding and hassling reviewers.
It's often not very good: an overzealous typesetter replaced every single italic p (as in "p < 0.05") with a musical quarter-note. Still, some of that money must be going somewhere.
That said, I would actually love a journal that offered (good) copy/line editing of accepted papers so that the science is presented as clearly as possible. I suspect this would work out well for the authors and the journal. At the margin, easier to read => more cites. However, very few journals compete on "author-friendliness". It's all impact factor this and metrics that.
A nomogram got futzed around and, well, the ruler-only, no-calculator required, calculations were now wrong.
"A set of trial calculations indicate that the nomogram of Burton et al is accurate, but a realignment of Burton’s nomogram (for greater symmetry?) in the Clinical Nutrition text resulted in a nomogram with considerable inaccuracy."
They did and do this sort of editing, so it may be, sometimes, a matter of journal quality. I don't anything gets into JAMA without that level of editing.
Anyway, that's what my opinions in this thread have been based on, but maybe I'm exhibiting selection bias by having experience primarily with a higher quality publisher.
I certainly didn't get any suggestions aimed at improving the readability of the text; those mostly came from coauthors and colleagues before submission, with a few more things flagged, in passing, by peer reviewers.
That's on the biology side of things. For ML, it's basically "Here's the doc class! Good luck!"
Or look at the editing that Elife professional editors perform in addition to other peer reviewers: detailed feedback on the manuscript for the authors, including requests for revisions and suggestions for improvement." detailed feedback on the manuscript for the authors, including requests for revisions and suggestions for improvement
Maybe poor quality journals don't do this sort of thing, I'm not familiar with them & their economic models.
They (the journal owners) don’t do these anymore:
-Proofreading: you’re expected to do this before you submit, and the reviewers will reject it if it is of poor enough writing quality. They will not, however, edit the paper in any way.
-Copyediting: there shouldn’t be copy editing in research papers, if there is fluff there is wasted real estate
-Typesetting: once again, researchers are expected to submit a paper with the typesetting requirements fulfilled (or in a format which can be automatically adjusted with a template (e.g. LaTeX))
-logistics: not an issue anymore, almost all access is digital and printed by the reader as necessary (if the portal is not DRMd to prevent this)
Copy editing is about a lot more than fluff. Accurately & skillfully communicating research and results does not completely overlap the Venn diagram skill set of "excellent researcher". Research skills do not equate to communication skills.
Typesetting is a little more complex than character format, but you're right, especially in these days with better tools, it's not nearly as labor intensive as it used to be, though still absolutely necessary if the journal is physically printed. Much much simpler if it's online only.
Many journals are still printed and archived especially by libraries, so those logistics are still a factor, and there is a fixed amount of labor required if it's 100 copies are 10,000. After the 100, the required work is incremental.
Digital absolutely has logistics involved as well: building and maintaining a platform isn't free, BUT:
I would wholeheartedly agree that the financial models here are broken and outrageous. Even considering the publishing work I mentioned, whether we disagree on the scope or not, better tools have made that work significantly easier. Scalable digital platforms are also much more of a known quantity and cheaper: inexpensive turn key solutions exist that could be scaled with fairly mundane tech skills. In every way I can think of, the costs associated with turning a raw article fresh from the authors into a polished publication should be much much cheaper.
Accurately & skillfully communicating research and results does not completely overlap the Venn diagram skill set of "excellent researcher". Research skills do not equate to communication skills.
Postdocs are no more likely to have their Venn diagrams overlap significantly.
But in any case, journals don't do any semantic editing on articles, so the whole point is moot: anyone can be an editor as long as they can typeset and maybe spot typos there and there, and neither communication nor research skill is required for that.
a journal will never edit the text beyond typos. their editorials are written from other peers for free. What copy editing?
Certain things would get edited automatically without signoff, but only if they were needed to conform to the style guide: some of that was set by the publisher, some by the journal editor.
Take two scientific researcher of the same discipline & scientific skill level, and one of them might be a very poor communicator. Multiple levels of the editorial process addressed those issues, copy editing included.
Changes were not made without the author's approval: They were edited, proofs sent to the author, who signed off, or declined some changes while adding others. It would then be reviewed by the publisher's editors again, and if all went well a final proof was sent out for a final sign off.
I'm not sure why you doubt this sort of thing: I lived in that world, on the tech side of a small/medium scientific publisher, and saw it all happen & built some of the systems that facilitated the process. It wasn't wasted effort: margins weren't very big and where I worked they would have been happy to get rid of a few staffers if the process didn't require them.
If you published an article through a reasonably sized publisher in a somewhat well-known, did you different experience? If so, then either the practices vary greatly among different publishers, or practices have changed a lot since the early 00's (which may very well be the case)
It also depends on the author guidelines of each journal. Nature group e.g. has requirements about when your subscripts should be in italics. And i think sometimes they prettify images. I don't care about those things, and it doesnt change a thing about the impact of 99% of papers. Most authors submit great quality manuscripts anyway, judging from what i 've reviewed. In fact i m pretty sure i m doing more copyediting and suggestions while reviewing than any of the paid editors.
Another thing is, a ton of effort is wasted in reformatting (and sometimes rewriting) a paper to resubmit to another journal. Publishing is definitely a wasteful process.
There are always publishing costs, but it would be cheaper to set up a public utility for proofreading and uploading papers rather than the current situation
As an example from Elife, in addition to the peer review feedback process, Elife describes their editorial output as providing: "detailed feedback on the manuscript for the authors, including requests for revisions and suggestions for improvement." detailed feedback on the manuscript for the authors, including requests for revisions and suggestions for improvement"
In terms of the mechanics of publishing, it's all very little different than a traditional publisher, except for the funding model. In terms of public access, it is far superior: Publication costs are covered without exorbitant subscription fees and the public gets the benefits of the research for no cost of their own.
This is why I would like to see a move to grant models where the grant includes the cost of publication, and a requirement for the portion to be used to get knowledge into the public even if the results did not confirm the experimental hypothesis.
As an n=1, I had a professor run out of money from their grant, so he gave a sob story to the journal and they agreed to make it open-access for free. I guess it depends on the journal.
I guess he would have paid out-of-pocket if the journal didn't accept the sob story.
Starting in 2013, all US-funded work needs to end up publicly accessible within a year of publication. There are slightly different ways, depending on the funder, but it's the rule.
Canada's Tri-Council Agencies have had a similar policy from 2016 onward, ditto EU....
Older papers are tricky though. I'm not sure the government can compel you to make something available.
I hadn’t heard about this, is there a keyword I can use to learn more?
This is a "meta-policy" of sorts, in that it directs grant-making agencies to form policies of their own. A lot was written about it at the time, most of which includes its title "Increasing Access to the Results of Federally Funded Scientific Research."
It came up again in 2020 when OSTP considered revising the rules (e.g., https://www.sciencemag.org/news/2020/05/will-trump-white-hou...)
Depending on your interests, you could then look at individual agencies' policies.
- Here's the NIH policy: https://publicaccess.nih.gov Search for something in PubMed and there will be a "PMC Free Full Text" icon if it's available from PubMed Central. Some publishers just lower the paywall, in which case it says "free from publisher".
- An FAQ for the NSF: https://www.nsf.gov/pubs/2018/nsf18041/nsf18041.jsp
- The DoD hosts its own repository called DTIC. Heres' the policy stuff: https://discover.dtic.mil/policy-memoranda/ Open access is the last section. You can search at the top-right for DoD funded research.
- Not to be outdone on the acronym front, the DOE has a site called PAGES, which ostensibly stands for Public Access Gateway for Energy and Science: https://www.osti.gov/public-access
- Those are the biggies, but there are a ton of federal agencies covered by the policy. Here's the USDA's plan: https://www.usda.gov/sites/default/files/documents/USDA-Publ...
For me, the rise of preprints has blunted the impact of that exclusivity period: something approximating the full text is usually available, whether from the publishers, PMC, or bioarxiv. This might depend on a lot of discipline--neuro dove head-first into preprints and I think they're less common in other fields.
It's not obvious to me what the "right" fee is. The NPG fees seem bonkers. The PLoS journals and eLife aren't trying to turn a profit, are a bit lower but still in the few thousand dollar range, so there must be some non-obvious, non-trivial expenses.
Scihub is at least levelling the playing field and forcing the conversation to happen.
Bioarxiv is seven years old, and when it started, it was tiny niche service with a few dozen new papers each month. People were skeptical of the contents of preprints and journals were reluctant to accept (or cite) work posted there. Now, it gets thousands of submissions a month and is a de rigueur part of many researchers' (and some journals') workflow.
In parallel, governments and funders have put some weight behind open access. An open version needs to be available within a year for all US-funded work from 2013 onward; Canada's similar tri-council policy took effect in 2016(?). It's true that there aren't draconian penalties, but there are some small nudges towards compliance. The NIH wants a PubMed Central ID for any paper you're claiming credit for on a grant. Google Scholar nags about non open-access papers on your profile. Many journals (even initially-closed ones) even handle the whole open access deposition thing for you now.
As a result, I find I need the library VPN a lot less these days. On the other hand, I agree that supplanting the publishers is tough. A Cell, Science, or Nature paper still carries a ton of cachet and its associated career impact; ELife isn't quite there yet.
Then, Elsevier, Springer, SAGE , et al. should spend their money making ToC / abstract lists of the best of the best papers in the different subjects.
That´s the main value of these editorials, they should super specialize in that and leave the distribution/editing to others (they kind of already do, they just want to have control).
I may not want to pay ANY money once I am looking for a specific article that I know it exists. But I will pay good money for a "push" type of listing/abstract and maybe analysis of the best articles in certain subjects (its kind of what Feedly does but without any analysis/editorial but just dumb grouping).
It seems you are claiming that the future isn't making all scientific content available, period - but rather only that content whose copyright holders have decided to make available.
I whole-heartedly disagree. We must not submit to arbitrary restrictions on the copying of information; and we certainly cannot and should not wait for Elsevier, Springer, IEEE et alia to grace us with access to articles.
Also - if "Open Access" means authors have to pay a large wad of money to have their papers published - that's not tolerable either.
Sci-hub existence is the proof by itself that you don't need billions to publish millions of articles.
While it keeps access open, it costs ludicrous amounts to submit an article (up to 4k $).
Not to mention the many journals focusing on OA that keep popping up who are purely profit driven, and care little for quality and standards.
With the severe lack of funding in the academic world, OA is not going to be the future, but will potentially skew access and impact towards research with deep pockets (rather than quality).
/side note: my spouse is an active publicising academic.
Cryptocoin community: hosting SciHub should be your platform's Litmus Test. If you can't do this one thing, your anonymous, decentralised, anti-censorship platform is a scam, so GTFO.
This is the kind of thing where getting a scuzzy blockchain involved would be ludicrous step backwards. Torrent isn't new tech but everyone seems to have forgotten it exists.
Incentivized onion routing networks like the Oxen network or Nym offer the potential for fast, anonymous filesharing, while decentralized storage networks like Sia or Filecoin could act as censorship resistant repositories.
Of course, the persons who make a clone will have to be brave, because they will face the same problems that Alexandra Elbakyan has faced.
Q: Do authors actually _have_ the right to republish the final published version from the journal they submitted their work to?
Q: Where is an author supposed to _obtain_ the final, "official" journal-approved PDFs in order to republish them?
Unless I head for sci-hub, I don't have any of mine :(
The big publishers have worked very hard to step this being straightforward :(
> You really don’t have PDFs of your own papers?
Absolutely not. Where would I have _legally_ have got them from? (This was 25 years ago) For my first paper, I have a copy of the paper journal for that month, kept as a souvenir.
> Did you leave academia
Yes
> and delete your data?
No!
That makes sense. Sorry if I was a bit incredulous ;-) I've been publishing for about 10 years, so I'm only talking about the current situation. Things might have been quite different before academic publishing was all online, before arXiv, and before having a "personal website" was a thing. If the publishing agreements back in the day already had a clause allowing the author to privately redistribute their paper, I would assume that if you as the author can get the published PDF anywhere (from a former colleague that still has academic access, directly from the publisher, or potentially even from SciHub), you should be free to then forward it to others or put it on your personal website. IANAL, obviously, so don't quote me on whether it violates copyright to get your own papers off SciHub ;-)
This plea to authors to change their behavior ignores why they submit papers to paywall publishers: The prestigious journal's acceptance of their paper helps promote their career.
Academic publishing is not a web host server for pdf files type of problem. Therefore, suggesting authors to upload their pdf to a public Dropbox url, GoogleDrive, Github repo, or their university faculty homepage doesn't solve the real problem. So even if Scihub had a "direct upload pdf" option, that still doesn't solve the underlying problem for getting their paper recognized for good work which spurs citations.
Scihub is a distribution mechanism for pdfs but not a recognition and impact filter for which papers are _important_. This is why scientists keep doing contradictory behaviors: On the one hand, they praise Scihub because it gives them access to papers -- but on the other hand, they keep submitting to paywalled journals to help their career.
Think of game theory incentives instead of hosting pdf files. Journals have the respected editorial staffs to look at their submitted paper and forward it to other peers for review. OpenAccess is a possible option but most OA journals don't have same prestige as the paywalled journals. That may change but it will take a long time.
An interesting question I think, is what value add does Sci-Hub provide, because obviously it wasn’t happening before Alexandra made it happen, does it outgrow her or is she holding it together?
Probably better to create a system that includes a reputation currency, since you can't 'spend' your reputation: https://medium.com/metacurrency-project/reputation-is-orthog...
That there are gigantic torrents available is almost useless. A popcorn time type of GUI client is needed that allows search, can dl the right chunks reasonably fast, seeds the rare chunks, has tit for tad implemented for a group of torrents. Could even do full text search by downloading all possible candidates after applying some bloomfilter.
I’m just in a “be careful what you wish for” state of mind, if there was no one in charge of sci-hub, the publishers could go on attack and fill the database with noise, copies of papers with numbers and methods altered.
Keybase makes it fairly easy, but people still have to learn what it means and why it’s trustworthy - my bet is things don’t change, there is a small percentage of people who understand how to verify the source, and the general population who either believes anything or nothing.
In the context of scientists and professionals tho maybe it is achievable to do some outreach and get people on eg keybase, something user friendly
I found it well explained etc, but it's hard to reach everybody.
That sounds what's needed, an extension of your IRL identity with top notch tooling to manage signatures (what PGP/GPG should've been) instead of docusign and similar that are charging an arm and a leg to "sign" PDFs with generated initials and signatures. We could go further; the signature is human readable in the rendered PDF (or whatever), and the Internet Archive takes a snapshot as an oracle notary when published, similar to Certificate Transparency for certs, but for all sorts of digital artifacts (the Internet Archive's snapshots have been used by a court of law for evidence authentication, so it's a trusted system [1]).
It ain't hard, it's just work.
[1] https://www.theregister.com/2018/09/04/wayback_machine_legit...
Digital signatures are actually overkill. The same infrastructure one uses to discover the authorship of a paper can distribute file hashes without any loss of autentiction, and I don't think anybody needs the non-refutation feature they bring. Maybe there is a nice use case for a paper authorship database with a PKI, but to my view the idea goes against an open scientific community that I think is much more valuable. (But again, maybe there is a way to have both.)
These are not basic problems. We're talking about scientific papers, not deeds to a piece of land or something.
> how do I know the pdf I’m downloading is what was published?
How do you know the paper photo-copy you have of an article is what was published? You don't 100% know, you make a reasonable assumption.
The exceptional case of needing integrity verification can have a niche solution.
Q: How open is the DOI system itself?
Long as the seeders stick around, this is decentralized, simple, and very hard to block.
Pretty sure this already exists, too, though it's extremely unpolished.
There are some projects that try to automate that, but so far all I have used needed some manual intervention.