Sci-Hub statistics and database
sci-hub.ru
sci-hub.ru
https://mobile.twitter.com/ringo_ring/status/143435621720862...
Scihub used to be a great resource, now it's only a resource for old research. Still useful for background material, but not for current work.
I also don't understand why the Indian court case has any impact on new article availability. The owner is not Indian. The servers and domains are not Indian. There doesn't seem to be any actual reason to stop adding new articles, other than some idiotic halfbaked point that only hurts the people who need the articles, like when Project Gutenberg banned anyone from a German IP, except this is much worse since there is no way around it for people who need new papers.
Because Sci-Hub has a good chance of winning the case. The court in question has previously backed a very broad definition of what constitutes fair dealing.
https://en.m.wikipedia.org/wiki/University_of_Oxford_v._Rame...
I understand that this is the party line that is parroted whenever this issue comes up, but it does not make any sense as a rationale for keeping new articles off the site. How is not adding any new articles (but, for example, keeping old articles accessible) assisting the possible winning of the case? And more to the point, why does it matter at all if it wins or loses the case? As stated, neither the owner or the infrastructure is Indian, so of what relevancy is this jurisdiction?
And further still, the case appears to have been delayed indefinitely. That last update claims that there was going to be an update a few days ago, but there was not. The proceedings are just now a list of one postponement after another [1]. Given that new articles are being held hostage, it thus very obviously benefits the legal system and the prosecution to continue to delay the case indefinitely.
[1] https://delhihighcourt.nic.in/dhc_case_status_oj_list.asp?pn...
As per the official tweet that has already been mentioned in this thread [1]:
> how about the lawsuit in India you may ask: our lawyers say that restriction is expired already
So according to the owner's official Twitter, this is no longer a valid reason, and yet new papers are still not accessible. Why is that?
[1] https://mobile.twitter.com/ringo_ring/status/143435621720862...
Give man a fish and he will praise you for a day. Give man a spinning and he will bitch at you because it is not a fish.
Guys: post your papers on the arXiv.
The court case has also been delayed for over a year now, so if it is delayed indefinitely, like it seems to be, then we will also not get access to new articles, also indefinitely? That's ridiculous. The last update from the court proceedings claimed that there would be a new update over a month ago, which in turn got delayed yet again to a few days ago, and there's been nothing [1].
[1] https://delhihighcourt.nic.in/dhc_case_status_oj_list.asp?pn...
Relatedly, if any academics want to help a small bit to resolve this by signing the amicus brief (i.e. intervention application) we made for Sci-Hub, you can do so by contacting me through https://docs.google.com/forms/d/1_if6Lipu-YPBMLk6zYjBxDFRA_c... and I will connect you with the coordinating lawyer. You can read more about this at https://forum.effectivealtruism.org/posts/bEKwqNDGysnZRcmpw/...
I am not sure how Sci-hub can get past this, unless they get a good Indian court ruling and can use Indian friends to scan printed copies of journals - if they exist in this online age?
I was struck by very low number of German downloads. Did I miss something?
Edit: also the sysadmin that keeps this database safe without 'accidental' data loss on UUID to downloader mappings.
I think it's very arguable that they have a legitimate interest here. Privacy has always been a weighing of interests, at least that's how I've always heard it explained by the Netherlands' face of digital law (Arnoud Engelfriet) also back in the days of WBP (the law from ~1995 that is 97% the same thing as GDPR), also in light of the European Convention of Human Rights (article 8 is a right to privacy).
A common example is filming the road: illegal, but if you park your car in front of your house and there have been car fires in your neighborhood lately, then it can be justified.
Filming employees inside a warehouse: invasion of privacy (illegal) but if there have recently been thefts from a certain part of the building then it's justified to hang up a camera there, introduce a lock that registers who went there at what time, or some such. (With adequate security measures so only authorized people can use it for the intended purpose.)
Personal example: monitoring everything I do on the company network is illegal, but because I work in a business where secrecy is important (security consultancy) it was considered justified to do spot checks, tell every employee upon entering into the employment contract that spot checks are a thing, and inform the subjects of spot checks after they were part of one. Transparent but still effective.
The two things to consider (iirc) are:
- Do my rights weigh heavier than the other party's right to privacy? (e.g. car fire is a fairly big impact on your right to the peaceful enjoyment of his possessions)
- Is there any other way in which I could achieve this goal with a lesser impact on the right to privacy?
In the case of Elsevier, from what I heard this whole scheme is a big mafia-like practice (wouldn't want to be published in a niche corner nobody reads now would you?) and so in my opinion it's entirely unethical to support (work for) them in the first place, at least in any role except one where you think you might be able to nudge things in the right direction. But I could see how a judge says: well, that's how today's law works, that you have moral objections is something you can take to your favorite religious leader and lament about, not a court of law.
If I'm being fair, there isn't even really an invasion of privacy because PDFs don't have executable code (usually) that can track you. Rather, they need to hide it somewhere so that, if it appears on the pirate bay, they can read out the ID and see who the perpetrator is. More like a criminal investigation using a fingerprint on a glass, and less like a cookie actively sent with every action you perform on a website.
TL;DR: GDPR applies, but it probably doesn't make this database illegal. It's not a loophole by which a person can say no to literally everything. (Would be cool if you could require the police to stop using your fingerprint in a legitimate investigation.)
Still, if I were that sysadmin... I probably wouldn't 'drop table elsevier', but I'd rather live off government benefits than support that scheme.
Quite trivially, actually, thanks to the good old analogue hole: https://news.ycombinator.com/item?id=30084193
All sci-hub would have to do in this case is download the same paper through three or more accounts (different institutions, networks, countries?) at three different times, rasterize them and keep the common denominator. If a pixel has no common denominator, they'd have to fall back to a default value. This is by no means a perfect method and it has its weaknesses and pitfalls, but the resulting PDF will be far less useful as a means to de-anonymize accounts using information from the PDF itself.
Publishers still have other sources of information to de-anonymize accounts if the multiple accounts/downloads aren't truly isolated from one another.
Or is there some other reason for the discrepancy?
Plus I would imagine that everyone wants their illness looked into, thus that's where funding tends to go. I care much less for physicists to figure out what dark matter is than how to treat health problem X that bothers me daily. (Just an example. In my particular case I'm healthy and would actually be quite excited about dark matter findings compared to any individual illness solution... but still.)
This comment would be worth a lot more with some stats about funding going towards the different fields, though. Not sure where to find that.
There are several medical subfields that on their own get more funding than the entire NSF (i.e. all other science).
Relevant xkcd - https://xkcd.com/2085/
Do note though that most math and ML practitioners use arxiv over sci-hub.
I don't know if sci hub bothers with publications that are available freely from an official source.
Don't use the term "open access" like this. A paper published on arXiv is free to read, and was freely published. "Open access" is a scam by the big publishers, where they don't take money from the readers, but make the authors pay. Or, putting it another way, anyone can pay their way in those journals and publish (sometimes sub par) papers.
"gold open access" is where you publish to a peer-reviewed open-access journal, which may or may not involve the author paying for the privilege.
"green open access" is where you publish to any peer-reviewed journal, and then the author self-archives the paper somewhere, like an institutional website, arXiv (as a "post-print", not a pre-print), or even Sci-Hub.
There are discussions involved about copyright and license and so on, but that's the gist of it.
Either way, we don't pay anyone any fee to publish on arxiv.
"arXiv (pronounced "archive"—the X represents the Greek letter chi [χ])[1] is an open-access repository of electronic preprints and postprints[...] "
https://openaccess.mpg.de/Berlin-Declaration
Publishers do misuse the concept though. They try to stay clear of using the term when they do. They use terms like Free Access or some of the more dubious colour variants of OA that doesn't provide all the freedoms that the Berlin Declaration of OA defines.
Interestingly articles uploaded to arXiv with the arXiv.org perpetual non-exclusive license are not OA as the reader is not allowed to redistribute the paper.
This is not true and comments like this are damaging to science.
Open Access papers are still peer-reviewed and by far not all of them make it into the journal. You can't pay your way into those journals.
Of course there are shady pseudo-journals which just cash in on the fee, do not carry out peer review and just dump the paper on the internet. But any scientist should be able to tell such scam journals from serious ones.
True, some journals, like many of the Frontiers series or PLOS One, make it very hard to be rejected in peer review. As long as your paper is reasonably well written and doesn't contain falsehoods it will almost certainly make it to publication. Still, you don't "pay your way into those journals".
Granted, many of papers in those journals report mere incremental progress. But these journals are still attractive for scientists to publish in, for obvious reasons. Publishers like PLOS use those journals as cash cows to fund their higher-tier offerings.
It is fair that the author pays for publication in those journals, since the most benefit is often for the author, not the reader. For the progress of science these offerings are not so useful, unfortunately.
She has done more than any other organization or individual in the history of mankind when helping people in second and third world country pursue advance research since the advent of internet. Well she and the people who pirate and distribute MS Office. Faculties around the world recommend scihub as the main and only source of research and journals.
It is not. A large batch of new papers was added manually, but the old service of typing in a DOI and having a paper be retrieved automatically is not working. Pick 10 random DOIs from 2022 and see how many Scihub will return.
This has been ongoing for a while now:
Rescue Mission for Sci-Hub and Open Science: We are the library https://www.reddit.com/r/DataHoarder/comments/nc27fv/rescue_...
Sites like ResearchGate make this very easy. And often a simple email does the job, too.
Advantages:
* It is legal
* The author gets feedback that someone out there reads their research
* Making direct contact to your peers is a good thing.
A community of scientists sharing their papers would be a good thing already now.
I personally know active scientists who don't even try anymore to look up the paper, but rather go directly to sci-hub for any doi they need. I can understand why, but I also think that this doesn't lead to a sustainable publishing culture.
Are you sure that the usual suspects don't make authors assign copyright or at least distribution rights? I wouldn't put it past them...
But honestly, how often does one have to skim that many papers in a day, to a level where the freely available abstract is not sufficient?
Perhaps every once in a while when one compiles a survey of a new field they enter. Once the project is set on the rails, one rarely has to read that much.
More often than you might think.
To take an example from my own work, I was doing assay design a while back, and needed to collect all existing primer sets in the literature. I probably went through a hundred papers over a several day period.
why are our scientists made to rely on elsevier et al to sift through the junk and find for them the perfect paper instead of doing it themselves? is science now such a cutthroat quick competition that it requires you to give a company the priviledge to work for you so that you dont have to do your own due diligence?
in india, we have a lot of local research that is done on open databases like shodh ganga and many more. but if you have to access foreign research material, better luck your university has an agreement with elsevier and others to pay them millions for a login. the alternative, go to scihub and find what you need.
i understand the whole quality/delivery debate but doesnt the average user already know who the big players in the specific domain are and who are trusted? or you want discoverability at the hands of a "trusted third party" without doing the legwork yourself.
then at the other end you have non-academics like me. I might have heard of a research paper in some article and i cannot read it without paying an arm and a leg. why? if we use the whole ebook/book argument that compensation is commensurate to the sales so more popular book means more money to the author but here authors arent compensated but elsevier so why should i pay elsevier? because they filtered through 1000 papers to provide 10 and for that privilege, they require unlimited royalty for ever? why?
isnt there a need of change in the "social ideology" of such schools and this whole elitism would go away?
Scientists do do that themselves. That‘s why it is called peer review. Journals take scientists work for free, they just pre-select papers, but don‘t do the review.
is the pre-selection such an important thing that these companies necessarily have to remain in business?
That we have to rely on single person effort that's risking their freedom to do so is fucking insane.
Nobel is a political tool that's mostly there to make a point (especially that peace prize).
https://slate.com/technology/2022/01/ibm-watson-health-failu...
The best “EHR” data we have—quantitative and minimally biased—-are from large genetically diverse animal cohorts like the BXD mouse family.
Basically: medicine as a whole is already some sort of expert system.
- Data collection and cleanup: Researchers conduct experiments to produce meaningful data and extract conclusions from that data.
This part isn't more automated because we have strict rules that prevent medical data collection and analysis without a clear purpose. Otherwise we'd be able to collect a lot more information to try and extract results from it using more inference-oriented techniques (deep learning and the like).
- Modeling & training: Expert panels produce guidelines from the results of that research. These panels are the "training part" of the system.
As a sibling comment said, replacing these panels with ML-based techniques isn't trivial because the data produced in the previous step is fairly noisy (p-value hacking, difficulty of capturing all the variables, etc.). Furthermore, the techniques that yield best results nowadays also produce them without clear explanations on why they hold, which is not something we are prepare to accept in medicine.
- Execution: Doctors diagnose and treat following said guidelines. In fact, they use decision flows that they themselves call... algorithms!
The main reason why execution is not automated is that we do not have the technology for machines to capture the contextual and communication nuances that doctors pick up on. There can be a world of difference between the exact same statement given by two different patients or even the same patient in two different situations. Likewise, the effect of a doctors' statement can be quite literally the opposite depending on who the patient is and their state of mind. One of the most important aspects of the GP's job is to handle these differences to achieve the best possible outcomes for their patients.
All that being said, there are companies trying to produce expert systems to help doctors diagnose. See https://infermedica.com/product/infermedica-api for instance.
Alexandra Elbakyan, who created SciHub, was born is Kazakhstan (exUSSR) and later studied in Russia where she lives now. I'm proud that Russia is one of the few countries on the planet where American exterritorial law doesn't apply.