Introducing The Paper Bay
jacquesmattheij.com
jacquesmattheij.com
In many cases that I have come across, the DOI can be turned into a link that takes you to the article. Perhaps that would be enough.
This would be a very simple but extremely helpful change. You only need to prepend http://dx.doi.org/ in front of the DOI, and make the resulting link clickable.
edit: done.
Also, while it's okay to have a somewhat complexe form to fill when making a request, it's really sad that it's as much work to share a paper, it should be very easy, in particular when clicking "I have this, send it", the form should have only one field: the file upload, or at least have the other prefilled with the request's values.
Maybe it's different in mathematics, I don't know, I'm just stating my experiences. Few fields are as long-lasting as mathematics, too, I think; in most fields, only a hand full of papers is still relevant after 20 years.
Sure, for applied (or even - experimental) sciences it is different.
Who does theoretical physics / mathematics and doesn't have access to a university library? And don't say Ramanujan, people in those circumstances wouldn't have access to the internet, either.
Of course, 'knowing one person' is anecdotal / talk about bad representative samples, but I can attest to the fact that those people exist. (I don't know the details why that person in question cannot register under a public library; public libraries may not have subscriptions to their journals of interest perhaps.)
In any regard, there are cases where old theories / research were dug up by contemporary researchers who realized that those original models were useful for their modern research etc. [citation needed, could dig something up, but googling is probably more effective than asking in this case.]
Let me state again that I too would sometimes prefer one huge easily accessible database with all articles ever published, along with cites and H-factors and objective impact factor rankings and a bookmarking/personal library feature and maybe a pony too. Then again, the magnitude of that problem is miniscule compared to other problems I have I'd much rather see solved, or spend my time on. And my experience says most of the people in my work field feel the same.
Google Scholar?
> How can a paper be perennial if it doesn't have any cites?
It was a bit exaggerated. But _in mathematics_ typical span of citation accumulation is decades, not years. And typical total citation count is way lower than in, say, biology.
> A paper being perennial is, I'd argue, defined by getting cites on a regular basis even after many years.
No. http://www.thefreedictionary.com/perennial
Don't confuse it with "popular" or even "with lasting popularity".
Again, in mathematics things (almost) do not age...
Sure, I'll accept that. My point is that some people must have it apart from the archives of the university it was first published at. Again, I'm talking about the actual, practical issues here, not the "what might happen". Not to turn this into an ad hominem, but are you an academic? How often do you have real problems (as opposed to 'annoyed because I have to spend 15 minutes') finding the content of papers?
"No. http://www.thefreedictionary.com/perennial Don't confuse it with "popular" or even "with lasting popularity"."
By that definition, anything written is perennial. In the context of a book/movie/paper being 'perennial', 'perennial' means 'still after a relatively long amount of time enjoys some form of popularity or following'. Just because a dictionary doesn't define it into that nuance, doesn't make it not true.
You are not making a convincing argument that papers are already "free enough". Unless you assume all interested parties are from the academic world, that is...
But seriously, how many people aren't? Yes yes there is always that one guy in his attic or log cabin... Most university libraries offer subscriptions to externals, too (mine does for less than 50 USD a year). So it's really only for people who are interested in academic materials, who are outside of any reach of a university (because really all you have to do is become member of a university library to get access to the electronic materials from the comfort of your own home). How many people fit that criterium? I'm arguing that overall, the problem is mostly in people's minds.
Can you show me one that does? You're saying as a non-institutionally affiliated individual, I can pay a university library $50/year and get access to papers? Please, which school offers this, or what search terms should I be using? please. times a googleplexamonium.
So, if you, for example, went and paid the UNC-CH library your $25.00 fee for a "borrower's card"[1], you don't get the ability to access all the various digital databases and what-not from your home, but you can drive down to the library, use the computer there, and access basically everything anyone else can. Or you can go down and ask the reference librarian to hunt down a paper for you and get you a paper copy.
On a semi-related note... I'm not sure what other states and jurisdictions have something like this, but here in NC we have something called "NC Live"[2], which is a portal that provides access to all sorts of online digital resources (including many which would otherwise be fairly expensive) to anyone with a library card from pretty much any county/city library in the whole state.
The key problem with journals is that much of what they do has been made obsolete over the last couple of decades by the move to electronic document distribution and the wonders of the WWW. Don't forget we are talking about private companies restricting the publication of academic research much of which has been funded out of the public coffers. Even the peer review is done, for free by academics. All that is left is the management of the process and distribution of the publications both of which I'd wager could be handled by the communities themselves as beautifully and repeatedly demonstrated by open source software projects.
What the journals still retain is essentially kudos, but in a very real sense as academics are largely measured against their publication performance of which the 'currency' is the impact factor of the Journal. Academics need to publish in those Journals with a high impact factor so that their departments get a good rating and thus get more government funding (at least here in the UK). The key will be for a community lead Journal to attain a decent impact factor. Soon.
Your last two paragraphs illustrate my point, or rather the myopic viewpoint that these 'paper liberators' advocate - nobody outside of the HN crowd I know (and yes, this is limited to a few fields) even cares about this supposed 'stranglehold'. Well everybody grumbles every now and then, but at least nobody cares enough to really push for change, and the things people grumble about (slow editors, idiotic formatting requirements, ...) are mostly not solved by 'open access' journals. Everybody has access to their universities' libraries, or contacts the authors themselves, or has some research assistant dig them up for them. Yes yes the subscriptions cost the universities real money, but come on, much more money is wasted by universities on other trivial stuff. Look, I'm not saying the model is perfect, it's just that (IMO) it's not a big enough problem for actual researchers (as opposed to 'interwebs poasters with strong opinions about things they have little use for') to really matter.
This division between researchers and everyone else is exactly the kind of elitism that broader access to research is supposed to reduce. People fighting for open access aren't necessarily responding to an existing widespread demand, but anticipating the future benefits to society.
They are trying to be ahead of the curve, somewhat like the oft maligned RMS was ahead of the curve with his understanding of and fighting against DRM.
"breaking their contract with the publisher"
I don't believe this is true, generally. of course, it depends exactly what they email you, but if it is a pre-print, that is, the PDF you compiled with LaTeX on your laptop, then there is normally no problem.The publishers only own the copyright to the particular typesetting of the paper (which was originally the useful service they provided academics) and as such, without their special-sauce theme that adds the journal name and page numbers on, it's fine for you to distribute it, host it online for free, etc. (Incidentally, this is also the case with sheet music, which is why it's legal to enter a score into a MIDI program and print it out, but not to photocopy it.)
As I say, this is probably not a universal fact, and will depend on what the terms of the journal submission was, and indeed for the book mentioned there was probably a specific contract at some point. Although, if there were small changes between the book and the thesis, the author was probably within her rights to email (and indeed host online) the thesis for free.
And Google Scholar often turns up pre-prints, making them much easier to find:
Having a few academics e-mail an author to ask for a copy might be acceptable. Having thousands of educated laypeople constantly e-mailing authors simply doesn't scale.
Another reason is to permit anyone and everyone to do wide-scale analytics, instead of having to rely on the analyses that big, well-connected entities decide to run.
Have you tried to do active research without access to a university journal subscription? I have to ssh into an on-campus box at least two three times per day to get a paper. I won't necessarily read the whole thing, but glancing at it certainly helps me see what it is about much more than a one paragraph summary.
If I didn't have that I would be much less productive.
The idea of using Tor is that the site is way harder for authorities to get down, and even if they do break Tor and get to the server, the users will still be entirely protected (there can't be any information about them in the logs or anything).
link: http://articleak.allalla.com/
submission: http://news.ycombinator.com/item?id=5175234
Moreover, a web1.0 interface does not seems to be efficient.
In that line there is already: http://www.pirateuniversity.org/
And... judging by it's name (i.e. The Paper Bay) I would expect something much more radical, actually aiming at getting 75TB of _any_ paper content...
So, a reliable, automated and anonymous (?) way to upload books/papers seems to be a must. Or do you have other plans?
I was working at MWC when Aaron was born. I think I met Aaron once when he was two or three, just briefly. I knew him from stories told by others, and eventually by following his writings.
MWC was quite an interesting company in its own right.
I believe the emacs that you are talking about is MicroEmacs, written by Dave Conroy while working at MWC. The source has since been made available generally, and the editor is the one that Linus uses. Daniel Lawrence later took over the distribution and here is one location: http://www.aquest.com/emacs.htm
I used it for quite a while, but then fell into Emacs and didn't look back.
Oh--I did get it to run on my HP200lx and used it there for a while.
Ok. done.
But I'd love to have just seen the pirate bay duplicate itself and to create a site exclusively for academic papers. Seeding of obscure papers might be an issue, but I'm sure that it would work quite well.
> When you download a paper from JSTOR or ScienceDirect, it
> is quite easy for them to add various watermarks to the
> PDF, and they already do (mildly).
You can strip the watermarks with a tool called pdfparanoia.https://github.com/kanzure/pdfparanoia
disclaimer: I wrote pdfparanoia because watermarks.
I don't care if the guy smuggles cocaine on the side, this specific thing he's doing seems like a really useful service.
Tracking down a paper is time consuming enough, even if you are sitting on the network with access to most journals. Don't make it look like people have to type in additional information manually.
However, I would not be as hard as you seem to be with the pirate would made The Paper Bay. When you really need a paper and you already lost half an hour searching for it in the web ocean, it's okay to take 2 more minutes instead of 30 seconds to ask for it on such a service.
Could there be a js solution to this? A volunteer inside the paywall could leave a tab open which periodically goes to get requests from the main paperbay queue and tries to fulfill them (you would need some robot-like functionality like find the PDF link). If success, it uploads the PDF. If fail (no subsc?), it can notify paperbay to put the request back in the queue.
With one requests every 10 minutes and 1000 volunteers you could have a 6000 paper/ hour rate of exodus.
BTW? Where are you hosted? How long do you think before they come for you if it gets big?
________________
PS: The protocol could be extended a bit it we need to handle captchas' as well. These really piss me off because I can't get the paper via ssh to campus + elinks!
R: requestor F: friend (inside paywall)
R-->tpb.com doi:10000x200 PLZ
tpb.com: resolve doi:10000x200, prepare scrape recipe.
tpb.com-->F could you get jrnl.com/yr/issue/33131/
F-->jrnl.com GET ...
F<--jrnl.com CAPTCHA.jpg + form el
R<--tpb.com<--F solve plz ( CAPTCHA.jpg , form el )
R-->tpb.com-->F form ans
F-->jrnl.com captcha form submit
R<--tpb.com<--F<--jrnl.com PAPER.pdfI wonder though how this ' we will deliver the paper on behalf of that email address.' will be important to challenge any copyright violation claims...
http://arxiv.org/help/api/index http://archive.org/about/faqs.php#Archive_BitTorrents
(And it just occurred to me that that's journalisms way of saying tl;dr. Except I did read it, just to find what the heck The Paper Bay is :)
There could be various settings to ensure each node with access to journals throttle the amount of downloads.
And I had no idea about Mark Williams C.
Relying on the benevolence of others may only go so far...
Of course the 'abuse angle' was taken into account, it's a bit annoying but any user facing service will have that aspect.
It'll never go down...
I don't know if such a website can survive to the scientific publishers' lobby, especially when its builder is publicly known.
Edit: it seems this exists, as another comment point out on this discussion page.