Detecting PhantomJS Based Visitors
engineering.shapesecurity.com
engineering.shapesecurity.com
I'm not sure as a website owner I want to only serve humans. What happens if someone bookmarks my site with a service like Delicious or Pocket or Pinboard and that site wants to scrape metadata from my site? Poor user can't see the metadata because I have a strange preference for human over robot users.
Oh, sure, those sites should get API keys, right? Err, no. Why should someone need to be able to parse arbitrary JSON and have to sign up with an API key for every site they visit?
Your web site shouldn't be (significantly) heavier or less stable. In an ideal world, if you followed REST properly, there'd be no real difference between an API and a website: both are just hypermedia representations of the resources you make available.
You can put data inside HTML (microformats, RDFa). You can use content negotiation so that you can have the same URLs serving up different content types.
Also, by filtering programmatic access, you'll piss off a lot of geeks who use stuff like PhantomJS to automate boring tedious shit that your website probably makes us do. I have little scripts that do stuff like automatically download invoices from suppliers for accounting purposes. All in order to prevent a security threat that shouldn't exist because you shouldn't have to rely on filtering particular browser types for your site to remain secure.
EDIT: seems more complex to remove than I thought: https://bugzilla.mozilla.org/show_bug.cgi?id=757726
[0] http://www.chrisle.me/2013/08/5-reasons-i-chose-selenium-ove...
[1] http://www.assertselenium.com/headless-testing/getting-start...
[2] http://blogs.adobe.com/security/2014/07/overview-of-behavior...
I'm not sure what the point of this is. An honest title would be "detect phantomjs by trusting the client". It's completely stupid.
The blog post on detecting PhantomJS is from Shape Security. The "fun fact" is that PhantomJS was developed by someone at the same company.
If you really want a web which only humans can browse and which prevents non-human-operated clients from visiting websites, require a submission of a blood, hair and urine sample or a photocopy of the person's driving license or something.
Until that point, computers will want to do useful things on behalf of humans, so it might be best to not get in their way because someone told you to on a website.
* Google wants to stop scraping so that you can't build a competing search engine that just scrapes Google for every search term that you see.
* LinkedIn has an amazing database of user information, you wouldn't want someone scraping all of it and creating LinkedIn2.
* One of the reasons Quora exists is that there are a lot of opportunities in mining the answers; they don't want another company to get to piggy-back on their hard work creating the site, acquiring users, making a good UI, paying for hosting etc.
It's still potentially onerous, since I regularly write agents that pull info from various places, for personal use, research, and archiving.