33GB of public domain JSTOR articles, and a manifesto
thepiratebay.org
thepiratebay.org
Process the pdf's with an OCR program to extract as much text from each document as possible. The extraction should be done page by page, so the extracted text can be referenced to a PDF page#.
Then, provide a searchable/browse-able directory of the extracted content. Each page of text has link to the original PDF page so you can easily open up the PDF to the page the text was extracted from.
I'd also make all text user editable wiki style. Combined with the inline PDF page references it would be super easy for any user to fix up translation errors from the OCR process. Tie in a karma system to the users profile so that edits can be thanked/kudos on a job well done to help with automating moderation of user edits by rating the user's current karma to decide if the edit should be accepted automatically or provided as an alternate version other users can check and rate up if they think it should replace current version.
Maybe mash in an image cropping service so diagrams can be cropped from the PDF and inserted inline with the translated text. Provide simple wiki formatting markup to allow users to format the articles.
Use ad revenue/donations to alleviate/cover hosting costs.
1, 2, 3, go.
Anyway the ethical thing to do here is to lobby jstor (a non-profit organization) to make reprints free.
So the guy has basically built a 33gb torrent of stuff that was already freely available to the public, just from a different source.
http://robocracy.org/royal-search/
https://github.com/bnewbold/royal-search
Lots of good old stuff, such as "On the Double Organs of Generation of the Lamprey, the Conger Eel, the Common Eel, the Barnacle, and Earth Worm, Which Impregnate Themselves; Though the Last from Copulating, Appear Mutually to Impregnate One Another"
This is all available elsewhere, i'm just trying to demonstrate how low the barrier to search is with OSS tools; I hope all these papers end up on archive.org with full text search!
http://www.pgdp.net/c/They have free journals in numerous fields, and gradually more big-name authors (in my field, neuroscience, at least) have been publishing in it. Its worth checking out.
Professors don't get much money from direct publication -- in fact, many conferences charge the professors who provide the content. They get paid by grants and their schools. Professors don't want their research to reach a limited audience. The universities doesn't want this either.
The only people in the chain who want limited access are the publishers, since this is how they make money. But between researchers, universities, grants, and publishers, the publishers contribute the least value by far. They generally rely entirely on other professors to edit their journals and conference proceedings, and for all the content. They charge ridiculous rates - often thousands of dollars for a single annual journal subscription - and get away with it because the system is not prone to change. Researchers are rewarded for publishing in "the best" journals, so no one wants to take the leap to publishing in a free space where there is currently much less prestige.
That's basically why academic publishing is messed up. Because there's money to made in keeping it messed up, and money to be lost in fixing it. But the ones who generate the real value _do_ want things to be as freely available as possible. If a critical mass of top-tier researchers agreed to stop publishing in non-free journals and conferences, it would probably start a revolution in this area -- but that's a lot to ask.
It's a prisoner's dilemma, in that the "traitor" researchers who keep publishing in the old journals will be rewarded.
(This is all about academia -- I guess motivations may be different in industry-backed research.)
http://blockexplorer.com/address/14csFEJHk3SYbkBmajyJ3ktpsd2...
Because they can.
When JSTOR was first getting off the ground 15 years ago, many of us had higher hopes for the organization's direction, that it'd eventually be a mixture of subscription-only access for recent non-open-access journals, and free-to-anyone digitization of old, public-domain materials, like a large-scale version of Project Perseus. But nobody at JSTOR really pushed the 2nd part; they digitized old documents, but kept them all subscription-only, and never even developed plans (despite occasional agitation) for individuals to buy subscription access for any sort of reasonable price. They seem to have developed a bit of an unhealthy attachment to "owning" the JSTOR archive as their valuable crown jewel, despite officially just being custodians of it on behalf of the academic community.
What follows is third hand information and not even directly related to JSTOR. So it could be totally wrong. But I am told that at the moment not all of the journals/articles in all areas are available to universities outside the US. Apparently it is hard to get hold of research related to subjects that the US thinks are not "exportable" because of potential impact on national security. Things like nuclear medicine, areas of chemical engineering etc. The journal publishers act as a point of enforcement in this story.
I guess that they are not a public company, therefore unfortunately perhaps we can not have any access to the revenue data to determine whether they are making any profit and how such profit is being used. In that case, most probably the safe bet is to decide that they are making profit. If, however, they are in fact not making any profit because their archiving and maintaining of these documents is paid off by the subscription charges, then this torrent does a disservice to science rather than acting as a liberator.
Without that information however I do not think it is easy, or indeed possivble, to rationally decide.
My guess is that their reticence to open things is more of an institutional megalomania of sorts. Sort of what you often find at art museums--- while their official mission is to promote the arts and educate the public, many of their staff are at least as interested in building the prestige/fame/fortune of their institution, which explains why they would have weird policies about reproducing their public-domain artworks.
Those costs must be made up somehow in the business model.
(before I get flamed, I'm not arguing that all documents should be $19 per article. I'm just saying that there's a non-zero cost that needs to be passed on to the customers somehow).
Sounds like a perfect job for bittorrent distribution....
Of course, a case could be made that university libraries should mirror all public domain digital documents of sufficient interest. Which shunts the cost onto the taxpayers.
Regardless, the cost still is non-zero, and someone has to bear it.
Likewise, archive.org does as well. I expect these papers will eventually end up in both these places.
Academic journals are full of incredibly specific knowledge. While some of it is accessible to people with a moderate amount of intelligence/diligence, any journal whose name isn't "Science" is designed for and by an incredibly small group of individuals with an incredibly specific knowledge base. Where is the donor base for that?
Answer: Universities are the only plausible donors. And they seem to think maintaining their place of privilege in being able to access the articles is worth some extra bucks.
Don't be so sure. Some experiments can produce large data sets, which ideally should be made available alongside the published paper.
But yes sharing with large academic datasets is a real and only partially solved (IOW not solved at all) problem.
But I think there is not a lot of interest in changing the system. The people who could change things are not directly affected by the disadvantages of the current system.
http://strategy.wikimedia.org/wiki/Strategic_Plan/Movement_P...
http://en.wikipedia.org/wiki/Wikipedia_talk:Lamest_edit_wars
A lot less cost than the current cost of paying subscriptions to that same content.
One issue that would need to be satisfactorily (and simply) dealt with is trust that the copy you have is the authoritative copy of record.
Academic publishing is a bitch and there is a lot going on lately towards a common, world wide reform. However, JSTOR is not really the bad guy here. Other publishers (e.g. Elsevier http://en.wikipedia.org/wiki/Elsevier) make billions by publishing mainly work payed with tax payer money.
Not in this case, apparently; they're public domain. Although JSTOR did foot the bill for the scanning.
I'll take your other point, though: if there's real evil in the academic publishing world, it's Elsevier.
I don't believe there's any part of copyright that takes the amount of effort into account. There was a case (again, I'll look it up when I get home) where a copy of a work that was in the public domain was granted its own copyright because it had a few errors in it. The errors were unintentional, but copyright was given. In the phone book case, the court specifically mentioned that even though there was no doubt a significant amount of work required to produce the phone book, that didn't give the publisher copyright.
Feist v. Rural
Can you clarify this a bit? How do you get to have impact without first publishing it somewhere (at least mildly reputable) and having other people actually read it/review it/run with it?
(On the other hand, these actions do contribute to create public awareness so I still haven't decided if they are good or not)
- http://lists.wikimedia.org/pipermail/wikien-l/2011-July/1092...
- http://www.generalist.org.uk/blog/2011/jstor-where-does-your...
Summary: they take in >53m USD. Publishers get 8m of that. Their IT infrastructure costs 4m. All the new material and scanning old material cost 5m.
(Notice how many millions are left over. They all go to staff, travel, and other goodies like that. The management is generously compensated.)
While there is great value in the JSTOR articles (all published before 1923)- in reality very few active researchers will be reading them. Some of the research will be just plain wrong, some corrected or superseded, and if there is any useful stuff left, then it will be readily found in much more modern and useful presentations.
I think the more important battle is to ensure all current research is deposited in open archives (preferably something like arxiv.org). The historical stuff will surely follow...
http://arxiv.org/help/support/2010_budget http://arxiv.org/help/support/whitepaper
It's hard to say how necessary a budget of that size is, made harder by the fact that they publish absolutely nothing about who they are, what they do with the money, how they plan for it, etc. The only source of financial information is their legally mandated annual IRS filing.
There's an attempt to sort some of it out here: http://www.generalist.org.uk/blog/2011/jstor-where-does-your...
Really? I think it's just politically less convenient than making laws that apply to future research. There's nothing that would keep the government from removing copyright all together if they wanted to.
1) It is unconstitutional to prosecute an individual for a crime that was legal at time of perpetration.
2) However, the opposite — decriminalizing previously illegal behavior — can be retroactive.
3) Changing future interpretation of copyright, etc., isn’t the same as case #1. If I’m not mistaken, Congress has passed e.g. the Mickey Mouse copyright law and the DMCA, which both extended “protection” & duration of copyright on previously created works.
It also seems we all benefit from allowing companies to invest in scanning public domain works since for whatever reason nobody is doing this by hand now.
So he clearly says its not in public domain but from a moral point of view they should be available to all the masses. The description of the torrent is very interesting. Definitely worth taking a look even if not willing to download the actual torrent.
I think this statement means that the documents are indeed in the public domain.
I also think equark makes a valid point about the scans. Further, I don't understand why these guys are going after JSTOR, a non-profit organization. I'd be more understanding of their methods if they went after somebody like Elsevier.
You have a valid point though. The access to these documents seem to be governed by agreements between various publishers with aim to share the published content among various institutions (http://en.wikipedia.org/wiki/JSTOR). So there is no point blaming JSTOR for the lack of access.
I do not understand one simle thing. JSTOR is not for profit. Are people therefore saying that it is making a profit? If yes, the action is understandable, but if not, as much as I am and was anoyed to face a JSTOR paywall and as much as there is an argument to make such paywall means tested, what are people actually saying? That this non profit organisation should not be able to existantially support itself?
Yes, government should support it rather than spend money on wars, but, I simply do not understand why anyone would want to kill off a non profit organisation.
How many of you are going to read these articles? How many of the general population would want to read these articles? Save for the professors and students how many would seek to even attempt to go through a journal article.
It just seems, frankly, that a self interested ideology of disdain for copyrights has sprung roots in Hacker News and man has stoped considering long term effects, focusing instead on their temorary pleasure of the now.
Whoever wants to. But doing so does not give you copyright over the content.
>I do not understand one simle thing. JSTOR is not for profit. Are people therefore saying that it is making a profit? If yes, the action is understandable, but if not, as much as I am and was anoyed to face a JSTOR paywall and as much as there is an argument to make such paywall means tested, what are people actually saying? That this non profit organisation should not be able to existantially support itself?
$19 per copy is extremely expensive, considering their expenses after the document is scanned are close to 0 (and as we can see, hundreds of people are willing to do it for free).
Bottom line is that people shouldn't make so much money off the back of unpaid academics/scholars. Personally I feel like if they charged something like 99 cents for a DRM free version of each paper (with free abstracts still of course) this might be more reasonable.
Just my take though :)
See: http://en.wikipedia.org/wiki/Bridgeman_Art_Library_v._Corel_...
Access to network resources is another matter--- Bing might be violating Google's access policies if they slurped Google Books, especially if they evaded rate limits or ignored robots.txt. But they wouldn't be in violation of any copyrights. This makes it a bit of a cat-out-of-the-bag situation. If I download a copy of an 1878 journal from JSTOR, and email it to a friend, I'm violating JSTOR's terms of service. But if my friend turns around and posts the PDF online, he is not breaking any laws.
A similar situation holds with U.S. government classified documents. Whoever leaked the Pentagon Papers broke the law, and could be prosecuted for the leak. But once they were out, as public-domain government documents, it was not illegal for third parties to republish them.
Of course, I don't begrudge people who scan the ability to get paid for their work. But there are a great many ways to get paid which do not involve forever robbing from the public domain (because, of course, if they are able they will simply keep updating their scanned copies while locking the public away from the originals).
I don't quite understand where the line is though. For instance, why isn't Google Map information public domain? The locations of roads aren't copyrighted and Google's efforts to enter this information don't seem that different than compiling a phonebook.
Most of the underlying information on Google Maps is public domain—road locations and physical features of the terrain are all uncopyrightable facts. What they claim copyright to are the map visuals: the particular style of drawing the maps, and the photographs they took. (though perhaps even that is debatable, since the photography is a robotic process)
You should be able to take the Google data and extract everything uncopyrightable from it to draw your own maps. (And if you're interested in public maps, you should see the OpenStreetMap project.)
But say you want to include some of the other things on Google Maps--place markers for hotels and restaurants, for example. Any one of them is an uncopyrightable fact: that's where it is. But what if you choose to mark the same collection of the same sorts of things that Google has marked--are you infringing on their copyright in the selection? The more you can make an argument that there was some original thought in the selection the harder it is to tell.
US law is pretty clear on plain facts being ineligible for copyright, even if it took work ("sweat of the brow") to compile them. But copyright in things like collections of data is much less clear.
Copyright law does not restrict copying of facts or ideas.
"(b) In no case does copyright protection for an original work of authorship extend to any idea, procedure, process, system, method of operation, concept, principle, or discovery, regardless of the form in which it is described, explained, illustrated, or embodied in such work." http://www.copyright.gov/title17/92chap1.html#102
Thus, you can freely copy the facts (information) in google maps.
Google may be able to claim copyright on map images it produces, but I'd expect courts would find that they are just a mechanical reproduction of factual information, limiting the copyright to just the presentation choices (colors, fonts, styling, etc.)
From the same reference: "(b) The copyright in a compilation or derivative work extends only to the material contributed by the author of such work, as distinguished from the preexisting material employed in the work, and does not imply any exclusive right in the preexisting material. The copyright in such work is independent of, and does not affect or enlarge the scope, duration, ownership, or subsistence of, any copyright protection in the preexisting material."
Several years ago I came into possession, through rather boring and lawful means, of a large collection of JSTOR documents.
But really, you don't need to be a genius to obtain large collections of scientific papers. Many people have them— its a prerequisite for doing meta-analysis for example. The distinction here is making them available to the general public.
I'm interested in hearing about any enjoyable discoveries or even useful applications which come of this archive.
- ---- Greg Maxwell - July 20th 2011 gmaxwell@gmail.com Bitcoin: 14csFEJHk3SYbkBmajyJ3ktpsd2TmwDEBb
Godspeed, good sir.
URL shorteners don't help.
I get a message about abusive content.
Some more info http://torrentfreak.com/facebook-blocks-all-pirate-bay-links...
Didn't want to write any blog posts about it, just point people to the original page.
Edit: Source: http://www.techdirt.com/articles/20090507/1152134782.shtml
http://www.google.com/search?q=577d58aa66beaceb71518ec417ab3...
The irony alone in the law profession would be tremendous. It'd be interesting to see if law firms would illegally access it: It would be obvious they were if they previously only subscribed to Westlaw, but the reality is most law firms subscribe to more than one database for emergency backup.
In my mind an easier way to disrupt this system would be to create a p2p site for article sharing - this often takes place informally anyway. Just a place where you could ask "Does anyone have article..." and then a friendly person would upload it to some filehoster and shares the link.
If not can someone post it so I don't have to get around the block.
[1]: http://cdmirror.textfiles.com/JSTOR_01_PhilTrans/1st_READ.tx...
For what it's worth, it would have been obvious they weren't facing a DoS.
"Unavailable anywhere else? Here's the ones from the 1600s: http://www.bodley.ox.ac.uk/cgi-bin/ilej/pbrowse.pl?item=ti... . Here's the ones from 1832-1938: http://catalogue.bnf.fr/servlet/RechercheEquation;jsession... . They're pretty widely available, for free."
"Looks like a bunch are on archive.org, too: http://www.archive.org/search.php?query=creator%3A%22Royal...
Can anyone confirm that these are the same articles in the torrent?
The only things in archive.org right now appear to be a half dozen issues from the mid-1800s.
Can you find this online? "Description of the Brain of Mr. Charles Babbage"
T1 - Description of the Brain of Mr. Charles Babbage, F.R.S JF - Philosophical Transactions of the Royal Society of London. Series B, Containing Papers of a Biological Character (1896-1934) VL - 200 SP - 117 EP - 131 PY - 1909/01/01/ UR - http://dx.doi.org/10.1098/rstb.1909.0003 M3 - doi:10.1098/rstb.1909.0003 AU - Horsley, V.
per wikipedia as of November 2, 2010, the database contained 1,289 journal titles in 20 collections representing 53 disciplines, and 303,294 individual journal issues, totaling over 38 million pages of text
http://www.transmissionbt.com/
Configure the bandwidth management to acceptable always on background levels and minimize to the dock or put on a different desktop with spaces.
Free art, technology and culture. Yep, this is the internet.
Now, a collection of all of Nature's issues would have been fascinating.
Thanks, gwern. Sorry for not RTFA first, but I've adopted a habit of reading comments first.
--- I believe these documents are still under copyright, so that would give them significant liability. They exist as a DMCA compliant company, so they are protected from their users' actions; but, pro-actively doing what you suggest would remove the DMCA protections.
> The documents are part of the shared heritage of all mankind, and are rightfully in the public domain, but they are not available freely. Instead the articles are available at $19 each--for one month's viewing, by one person, on one computer. It's a steal. From you.
Scribd makes money off ads, so the incentives are actually really well aligned