JSTOR Liberator sets public domain academic articles free
arstechnica.com
arstechnica.com
"At the same time, as one of the largest archives of scholarly literature in the world, we must be careful stewards of the information entrusted to us by the owners and creators of that content." -- JSTOR
The contention that JSTOR is acting as a faithful steward of public domain works -- having established any sort of restriction or barrier to prevent the public from accessing those works -- truly examines the widths and depths and bounds of intellectual dishonesty.
Let's see..
http://www.jstor.org/discover/10.2307/1831029
The "rights and permissions" link goes out to:
https://s100.copyright.com/AppDispatchServlet?author=Fischer...
... ok copyright.com is not encouraging at all, that scam is everywhere. Let's try something older.
http://www.jstor.org/discover/10.2307/30096268
The "rights and permissions" link goes out to:
https://www.copyright.com/openurl.do?sid=pd_ITHAKA&servi...
which leads to:
http://www.copyright.com/search.do?operation=detail&item...
.. where they seem to be selling the rights??
Isn't it better for a non-profit entity to digitize PD articles and make them available to the world at a reasonable price? Or would you prefer a world where JSTOR didn't exist and the only people who had access to most old PD documents where those who lived in major western cities near the biggest libraries?
In all seriousness, have you ever tried digitizing old journals? It is fairly time consuming.
Writing an encyclopedia is time consuming, but we did that.
Walking & driving all over the world is time consuming, but the OpenStreetMap community has done that.
Scanning & digitizing old public domain books is time consuming, but the Project Gutenberg project has done that.
What's next?
No, I have not tried to digitize old journals. But I HAVE helped contribute to the digitization and distribution of public domain works. I'm aware of the time and financial resource commitments involved.
I grant that a $19 article may represent subsidization in their system, but argue that they have an ethical obligation derived from their position as a steward to make the components of that price transparent.
It's my understanding that universities pay significant amounts for site licenses and that general public purchases do not subsidize that access.
By perpetuating a watermark or attribution on a public domain work digitized by JSTOR, are we as a society not implicitly granting JSTOR a right on the work which it ought not have? Digitization is a one-time salvage operation. Should we conflate it with copyright through watermarks?
Regarding JSTOR, journals, and non-public-domain works -- that is a private property matter. They ought to be able to charge what the market will bear.
> It's my understanding that universities pay significant amounts for site
> licenses and that general public purchases do not subsidize that access.
Yeah, I am pretty baffled with how non-institutional users are requested to pay for access, both on JSTOR and other publisher's sites. Surely this is not a large source of income? arXiv (always pointed out as a shining example of something going right) has recently started requesting membership fees to be paid by the top 200 institutional users. Other than that, the public is still able to go to the site and read about dark matter and substring bananas.You know, what probably happened was that when the content was starting to get digitized, the publishers only had users that were accessing from universities. So they charged the universities just like they previously did for subscriptions to the bound book issues. At the time, there was probably no concept of doing any sort of analytics to figure out where downloads were coming from, so they didn't think that maybe they should just allow anyone to download, and then setup a deal with colleges to pay their fair share. I think the arXiv model is worth some more investigation..
(This ignores all the actual truths like, they have to maximize shareholder value, they have to make as much money as possible, their incentives are aligned differently, etc. Instead it is simply focusing on the reason why "individual purchase" is an option that doesn't seem to make sense to non-institutional members like the general public.)
1) Some of the articles on JSTOR are the product of complete public funding. But many are not. Many are the result of partial or complete university funding.
2) All the work of editing the underlying journals, which is very time consuming, is generally not publicly funded.
3) All of the work of digitizing those journal articles is not free. When Google undertakes to digitize such content, they charge you by selling your privacy to others. JSTOR just asks you to pay a fee for their service. Who is the bad guy here?
Non-sequitur, and you can bet plenty of researchers will publish publicly in order to get free money. What it will affect is private industry benefitting from publicly funded research and resources.
The universities are not paid by the journals, — quite the opposite: they pay handsomely for access the work of their own researchers.
In most cases the journals are edited by the same people who write for them, for free (sometimes there'll be a pittance salary, but it's not enough to quit your day job by any means).
The "incomprehensible gibberish" is, in general, fully comprehensible by their intended audience, which is other academics working in the same field.
Here's how the model works:
1) Academic writes the article, supported by taxpayer funding, or a grant from a foundation, or whatever. 2) Academic submits the article to the journal, where it is peer-reviewed and edited by other academics. 3) None of these academics get paid a cent for their work (other than the aforementioned public funding). In specific, the journal publisher doesn't pay them anything. Some of them even CHARGE the author. 4) The journal publisher then sells the result for hundreds or even thousands of dollars per year.
Nice gig for the publishers. Less so for the academics and the taxpayers.
This model made a certain degree of sense back when journals had to be printed (short print runs are expensive, especially stuff with lots of diagrams, weird equations, foreign languages, etc. as journal articles tend to be), then physically distributed by paper mail to institutions all over the world.
That's no longer the case.
I know people who edit academic journals, and have published in them myself.
Your claim that editorial costs are what drives journal prices is completely without merit.
Few journals below the level of, say, Nature, have paid staff, and few academics would tolerate major rewrites of their work by an editor anyway.
The importance of copyediting a scientific paper
D. J. Bernstein
2005.05.04
http://cr.yp.to/bib/20050504-copyediting.txtAs for the point re: editing, this is just not true. Academic journal submissions are by highly educated PhDs at the best universities; the editors on the other side just aren't as well trained. Post-submission copy editing is minimal; most of the effort by journal editors is on formatting. The reason is that in many cases the journals have failed to invest in a proper publication workflow so they have to spend time converting Word documents into their in-house style, rather than just (a) distributing a LaTeX template or (b) working on a common interchangeable format and simple rich text editor. Some of the more computational journals, like Bioinformatics or NAR in compbio do have LaTeX formats, but many like Science, Nature, or Cell do not.
So their editorial costs for their human army arise because they didn't invest in a few engineers to build a simple rich text frontend with pdf preview functionality. They are manually doing simple asserts like checking character counts in headline text rather than doing that via client-side JS. Things like that are the cost centers here. A shame.
This is not true. Most editors are PhD-trained people (after all, they have to select the reviewers of a paper, which requires knowing the relevant people in a field who are capable of evaluating a paper). In some of the smaller, field-specific journals, editors are themselves volunteers from academia (go to your local university, find some lesser known journal and check the inside cover, which usually lists the editorial board of the journal. You'll see that most of them are university faculty).
Let me paint you a picture of a hypothetical alternative system: a publicly funded, 100% free-to-access online system where academics can submit papers, and then other academics can peer review them. Papers and authors are ranked by an open-source algorithm that takes into account factors such as the number of citations, the quality of the papers making the citations, results of peer reviews, the quality of the peer reviewers and so on. Doesn't that sound like a better place for our shared cultural knowledge, that belongs equally to everyone, than being locked-up and monetised by a handful of private companies?
But I did all of my journal editing as a grad student in college. Technically, it was my advisor's job to do it, but he was too busy with non-mundane things, so this kind of work got passed off to his students. None of us were paid for the work, of course. It was just understood that you do it because you have to (grad students need good recommendations, and advisors to those grad students need to play politics with the journal committee). As far as I can tell most editing for most journals is done for free by various university professors and (mostly) their students. More prestigious journals like Nature and Science probably have their own editors do an addition pass after this, but most don't seem to do that.
It might be different in some fields, but that's how it works in my field (mechanical engineering), and it seems to be that way in other engineering fields as well.
Yes, 34$ for a 20 page PDF. I could get printed books with collections of 15-20 hand picked articles for less. The obscenity is related to the huge & itemized price.
Also, if you are not a student at one of the affiliated universities (which is a rare case, probably less than 1% of the population) then you get no good options. How is that for advancing the arts and sciences?
How about we get all that has been funded by the public back to the public, and stop this madness.
Another problem is lack of access to the text for machine learning, NLP purposes.
In conclusion, I get they invested some money. They should just be nationalized. Pay them a compensation fee and just cut them out of the loop. They are an obstacle to progress.
It is one of those times when public good trumps individual property rights. Back in the time when they were building railroads in USA, it was necessary to solve a similar problem - how could they pass the railroad through the maze of public properties. The solution is simple - expropriation + fair compensation. Public good must be met first.
Here we are, nearly 40 years later, standing next to Wikipedia, OpenStreetMap and all the open source software in the world, and we can say, Yes, people will do this work for free. The fundamental premises of your argument (that people won't do it for free) has been disproved.
JSTOR's non-public-domain works are a different story, but the journals are the most to blame for that; JSTOR is just a licensee. Pretty similar to Google Books in that case, too (Google will only let you see a "restricted preview" or "snippet view", depending on the work, because they don't own the copyright).
That's not to say there can't be other scanning projects that aim to do a superior job, and I'd probably volunteer if there were a way I could be useful to such a project (I've spent some time at Distributed Proofreaders). But I don't see JSTOR as exceptionally evil, at least any more than Google Books is. Both are bringing more content online, in ways that are partly good and partly flawed.
It's absurd that copyrights are so lengthy that 1923 is the cutoff point, but that's a whole other can of worms.
Indeed, every time I see a publisher offering me an article from the 1920s or 1930s as a digital download for a small fee of around, oh, say, $30, my heart is filled with contempt for this attempt of milking every last cent out of every paper.
As I see it, they've removed the majority of the barrier (i.e. actually tracking down a copy of the periodical or whatever). So you have to log in and deal with a watermark- calling that 'intellectual dishonesty' is stretching the term to its limits.
Because a) they need ammo against people like, well, Aaron Swartz, and b) doesn't take many people bulk downloading the DDOS the sucker.
> Why does a one-time salvage operation require a watermark - are they asserting a copyright?
They probably don't distinguish between public domain and copyrighted in document production. TBH I doubt that most documents people access are public domain, so it probably hasn't been an issue until now. Even in fields that are really old (e.g. classics) the vast majority of work is recent and copyrighted.
Besides, if I were to scan these things in, I would want people to know I did it. Do you know how much effort goes into the process? Even when things are completely automated (which is expensive) there's a lot of manual labor in operating the machine, editing, and cleaning stuff up. What do you do with figures? How about typeset math? Glyphs not in unicode?
Not saying I agree with their tactics, but I understand them and I don't think it makes them an immoral/bad organization.
Step 1: Ask Google to do it. Step 2: There is no step two.
Google Books shows that Google is willing to do these kinds of tasks, as long as they can show the results on their site and thus keep people using Google services. I'm sure they would digitize all of the public domain papers and host them at no cost to JSTOR.
Does anyone chiming in who claims this info must be available for pennies per article actually have any evidence than an operation of this scale can be funded this cheaply? The costs to provide this service are not this simple and as cheap as you think.
Exactly. I can't believe bandwidth is jstor's limiting costs.
JSTOR sends messages into the marketplace that they are a faithful steward of the public domain, but knows that its TOS a) prevents unrestricted access to the public domain, and b) knows that a US Attorney will prosecute violations of that TOS as felonies requiring several years in prison; I argue that their speech does not match their actions and that this is dishonest. Aggravating this, the extreme negative consequences of taking them at their word (as Aaron did) is why I've selected to say that their actions probe the depths of intellectual dishonesty.
Full-text search, like printing, is a value-added service. There is no reason to keep public domain works hostage by the threat of sending people like Aaron to jail for decades just so that they can offer FTS. If JSTOR wants to offer FTS or other services on the corpus of public domain works, let them charge for access to those services.
An existence proof: arXiv offers bulk download access to the works in their repository (~490GiB) via Amazon S3 Requester-Pays buckets [1]. The requester is paying Amazon, not arXiv; so, arXiv doesn't earn "pennies per article", it earns nothing. Incidentally, arXiv also provides full-text search [3].
Let's talk about capacity planning. Let's guess that average size of a digitized journal article is 5 MiB, that comes out to ~361 TB, or $44,400 to stream from S3. The at-rest cost of those articles is far lower because there are far fewer articles than downloads (I don't have a number, but would you argue otherwise?).
My proposal: JSTOR removes its watermarks and puts all public domain works and associated metadata into S3 Requestor-Pays buckets. They finance their operations by selling non-public-domain works and value-added services like FTS at a price the market will bear.
Earnings made by restricting bulk access to public domain works is blood money. Watermarking these documents is confusing - no one seems to have answered my questions about whether they are claiming a new copyright.
[1] http://arxiv.org/help/bulk_data_s3 [2] https://forums.aws.amazon.com/ann.jspa?annID=386 [3] http://arxiv.org/find
How will JSTO be a "good faith steward" if they give all of their content away and make it available for free? They have to pay their bills to maintain the infrastructure to support all of this and pay their employees to do so. Charging for article access is how this is accomplished and how the company behind this allows the service to continue. There is a lot more money behind the operation that just copying pdf's to an s3 bucket.
Where we differ in opinion seems to be that I do not believe that a company can take a public domain work, wave a magic terms of service, and then sell it back to the public.
"Just because it is public domain doesn't mean it can be taken and posted elsewhere." No, that is precisely what it means - the collective commons owns this; ownership means having the right to use your property. These works are your property, my property, our property.
Time and money was spent organizing and publishing the book / article, so they do have a right to charge for it. If you want wholesale free access to public domain documents, then you find a place which will provide these to you for free. You don't get to ignore a companies right to charge for something just because you disagree with the premise.
https://groups.google.com/group/science-liberation-front
Someone contributed a small greasemonkey script that does something similar to the JSTOR memorial liberator, except for SpringerLink previews.
How hard is this logic to understand? If you don't want to use JSTOR don't use it. Don't go around saying they should let you download it for free.
Just because someone is selling something that is public domain doesn't mean that you can't give it away for free.
I'm sure she can rationalize that as a terrorist attack on essential infrastructure, so let's round it up to 170 years in jail for the conspirators. Let's assume every one of those connections would have paid $200 for an article,so that's $50 million in damages. Right?
The response time on an example PDF for me was around 10 seconds.
At this point, I don't think it would be a stretch to see an overreaction, even if it's just benign traffic from people checking out the free content melt the servers.
Accessing "hidden" URLs has been called unauthorized access; deep linking or just linking to something like deCSS has been likened to a crime; port scanning has been legally attacked a few times.
If these same type of people think URLs are a crime, I don't think the scenario I described is that bizarre. Look up that news story about the guy in a glider near a nuclear plant (he was catching a thermal from the lake) -- there was never a no fly zone, but officials actually considered shooting him down! Instead they held him for 24 hours while his loved ones started a search for his glider. They dropped the charges only after he agreed not to sue them!
My point is "Why don't we discuss things as they are, not as they may or may not be some time in the future?".
EDIT: Apparently the TOS (not EULA) is vague enough that this might actually violate the TOS- but the creators are very explicit about this and clearly tell the user that if they violate the TOS it's their own damn fault. The only thing I could think of is that the creators could be hit with DMCA violations for circumventing the 'DRM' of the site, EVEN if they never actually use it on anything copyrighted.
And you don't think that she will think that this is hacking if the site goes down from the load? Besides, making an example of people is having it out for them.
No, I don't. Someone stupid enough to continue this shitstorm would never make it that high up in office, especially when this is even less plausibly hacking than what Aaron did, even to someone without a great understanding (it's distributed, they warned people it would violate TOS, and the intent is clearly not to burden the site but to make a political statement that is explicitly designed NOT to burden the site).
I could be wrong. But I really, really don't think so.
> Besides, making an example of people is having it out for them.
Well yea, she might have it out for 'hackers', but who doesn't? Anyway, it's a common practice to go after the 'big fish' to scare the smaller ones.
Once I download the public domain article from 1886, I suppose I could do whatever I wanted with it.
Is scanning old journals a creative process? Maybe there is some case law that says so, but until I see that, I am inclined to say no, and everything that I can find agrees.
http://en.wikipedia.org/wiki/Database_right#United_States
http://en.wikipedia.org/wiki/Threshold_of_originality#Reprod...
http://en.wikipedia.org/wiki/Bridgeman_Art_Library_v._Corel_....
If JSTOR has a different opinion, they can see me in court.
So it's not as if JSTOR suddenly saw the light of day and opened up their archive for download.