American Chemical Society bans university after "spider-trap" is clicked
blogs.ch.cam.ac.uk
blogs.ch.cam.ac.uk
Atypon has [a relatively small client list](http://www.atypon.com/our-clients/featured-clients.php). Compare it to [Highwire](http://highwire.stanford.edu/lists/allsites.dtl). I'd be willing to bet that all journals hosted with Atypon share this spider trap—even journals that are supposed to be open access where spidering should be OK.
Scientific publishing is weird. Source: I work in scientific publishing.
That said, ACS is pretty sinister, too. They opposed PubChem (http://en.wikipedia.org/wiki/PubChem#ACS.27s_concerns) and in general don't behave like the nonprofit scientist trade organization that they present themselves as.
If you avoid caching requests by robots, then instead you end up having to go through all the layers of your app, possibly going in the database.
In most situation, I don't think the above matters much. But I can see why it could be a worst case for some stacks.
I mean, it's marked by comment tags that say "spider trap" right on them! Its the worst type of disambiguation system: likely to generate false positives, unlikely to catch real violators.
To really do damage, you need to create something that'll go viral, and piggyback the link on it. Undergrads will click anything that looks fun. Make it look fun and you're golden.
This was in the late 1990s.
Tl;dr: researcher is browsing source code of a research paper's web page and finds a strange link (but same domain). She clicks and is informed that her IP is banned for automated spidering.
Apparently, this research site is meant to be open-access...
-------
Pandora is a researcher (won’t say where, won’t say when). I don’t know her field – she may be a scientist or a librarian. She has been scanning the spreadsheet of the Open Access publications paid for by Wellcome Trust. It’s got 2200 papers that Wellcome has paid 3 million GBP for. For the sole reason to make them available to everyone in the world. She found a paper in the journal Biochemistry (that’s an American Chemical Society publication) and looked at http://pubs.acs.org/doi/abs/10.1021/bi300674e . She got that OK – looked to see if they could get the PDF - http://pubs.acs.org/doi/pdf/10.1021/bi300674e - yes that worked OK.
What else can we download? After all this is Open Access, isn’t it? And Wellcome have paid 666 GBP for this “hybrid” version (i.e. they get subscription income as well. So we aren’t going to break any laws…
The text contains various other links and our researcher follows some of them. Remember she’s a scientist and scientists are curious. It’s their job. She finds: <span id="hide"><a href="/doi/pdf/10.1046/9999-9999.99999"> <!-- Spider trap link --></a></span> Since it's a bioscience paper she assumes it's about spiders and how to trap them.
She clicks it. Pandora opens the box... Wham!
The whole university got cut off immediately from the whole of ACS publications. "Thank you", ACS
The ACS is stopping people spidering their site. EVEN FOR OPEN ACCESS. It wasn't a biological spider. It was a web trap based on the assumption that readers are, in some way, basically evil.. Now I have seen this message before. About 7 years ago one of my graduate students was browsing 20 publications from ACS to create a vocabulary. Suddenly we were cut off with this awful message. Dead. The whole of Cambridge University. I felt really awful.
I had committed a crime. And we hadn't done anything wrong. Nor has my correspondent. If you create Open Access publications you expect - even hope - that people will dig into them. So, ACS, remove your spider traps. We really are in Orwellian territory where the point of Publishers is to stop people reading science.
I think we are close to the tipping point where publishers have no value except to their shareholders and a sick, broken, vision of what academia is about.
UPDATE: See comment from Ross Mounce: The society (closed access) journal ‘Copeia’ also has these spider trap links in it’s HTML, e.g. on this contents page:http://www.asihcopeiaonline.org/toc/cope/2013/4
you can find
<span id="hide"><a href="/doi/pdf/10.1046/9999-9999.99999"> <!-- Spider trap link --></a></span>
I may have accidentally cut-off access for all at the Natural History Museum, London once when I innocently tried this link, out of curiosity. Why do publishers ‘booby-trap’ their websites? Don’t they know us researchers are an inquisitive bunch? I’d be very interested to read a PDF that has a 9999-9999.9999 DOI string if only to see what it contained – they can’t rationally justify cutting-off access to everyone, just because ONE person clicked an interesting link? PMR: Note - it's the SAME link as the ACS uses. So I surmise that both society's outsource their web pages to some third-party hackshop. Maybe 10.1046 is a universal anti-publisher.
PMR: It's incredibly irresponsible to leave spider traps in HTML. It's a human reaction to explore.
I think the case mentioned in the article is definitely heavy a heavy handed approach. When it comes down to it at my place we are just trying to block the wget -r's of the world.
Not sure it checks for styling before prefetching them.
In Chrome 32 there's a "Predict network actions to improve page load performance" config option, and here's what it does https://support.google.com/chrome/answer/1385029?hl=en
edit: and i was messing the webstats for advertisement.
We eventually figured out he his online album had an unprotected "delete this photo" endpoint via a GET, and no robots restriction! We eventually had to fix the crawler to detect things like this...
Looking at the source... there's some weird things going on, I think maybe the _original_ page loaded it's content with Javascript, and the google cached version is just the JS skeleton, waiting on trying to load JS from the original (overloaded) site which will actually load the content?
Ugh. The trend for JS-dependent sites for simple content breaks the web, people.
2. subscribe
3. click link
4. sue them for breach of contract and damages. (they didn't deliver the content you paid for, it damaged your main source of income: providing knowledge to paying students)
5. repeat.
There are bad actors out there, they exploit services, and one of the ways the services detect them is to create situations that a script would follow but that a human would not. When they do something bad you've got a couple of choices, cut them off or lie to them (some of the Bing markov generated search pages for robots are pretty fun))
So she sends an email to the address provided, they talk to her, she gets educated and they re-enable access. If it happens again the issue gets escalated. Its the circle of fraud.
I wonder sometimes if this sort of activity (honey pots for autobanning) angers people because they feel they have a right to script the site or if they feel poorly for "falling for" the honeypot. Clearly it generates some emotion though.
I definitely noticed that the tag was <span id="hide">, though... that raises all kinds of interesting questions like "what if I want to hide more than one thing on the page?" and "if I'm going to do this, why not display="none" instead of just adjusting the text color on the link?"
Lets wait and find out how long it takes them to respond to the inevitable interest that 999999.99999 people will have sent their way ..
There are several posts about this issue, http://blogs.ch.cam.ac.uk/pmr/2014/04/03/acsgate-the-america... appears to be the best I've looked at as it gives details of the links that were followed that initiated the suspension of service by ACS.
This goes against the nature of the internet and information, it is bound to be free.
This makes me furious. It isn't because the intent is malicious. That only makes me just a tiny bit angry. I am furious because the malice was implemented in the stupidest, most useless, laziest manner possible.
It's like keeping the neighborhood kids off your lawn by burying a pressure plate switch out there for the armed nuclear bomb in your garage. And then not telling anyone about it. And then inviting all the neighbors over for a croquet tournament.