Detecting PhantomJS-based visitors
engineering.shapesecurity.com
engineering.shapesecurity.com
Beware it's Affero GPL.
It implements most of the Selenium WebDriver APIs. Currently it's a work in progress especially in regard to persistent tracking. But it's capable of addressing many issues raised in the article. As a bonus, for Java users it's a lot faster than GhostDriver at least.
Header ordering: https://github.com/MachinePublishers/jBrowserDriver/blob/mas...
DOM objects (user-agent, navigator, Canvas, Date) are addressed by injecting the response content: https://github.com/MachinePublishers/jBrowserDriver/blob/mas...
JS Engine: about 1 year newer than PhantomJS. Java ships with Qt5 WebKit.
I'll have some significant updates in the next couple of days surrounding all these issues.
In my experience, these two use detection algorithms that could best be described as "pattern-based."
Hitting any of their pages once via any scraper will not get you blocked, but exhibiting request patterns that sufficiently differentiate you from human traffic definitely will get you temporarily blocked or forced to fill out a CAPTCHA.
For example, you can hit Google Search once an hour via a painfully obvious scraper no problem. But, even if you take control of a Chrome browser via selenium, and write your scraper to do everything exactly like humans down to typing in the searches and moving the mouse around and clicking results, Google is ridiculously good at identifying bot vs human traffic patterns.
I think the algorithm builds a "normal usage profile" for a combination of IP,cookie/user,device and sets a threshold, activity above that threshold gets flagged.
When working from within the network of a large company that uses a proxy I saw the CAPTCHA regularly for some time.
Of course the CAPTCHA gets answered mostly correctly in this case, which could trigger manual inspection and finally addition to the white list.
I once scraped Google Search a few thousand times, in a few seconds, a few times, just for fun, and in a school (1000+ people). It was quite funny when everyone suddenly saw a captcha.
Then a month ago or so I scraped Google+ for some statistics, maybe 1-3k times in 2-3 minutes and nothing happened.
The point is that my node.js scraper was far from perfect. Just a fake Chrome user agent, nothing else.
If a large organization with a static IP has a lot of users, Google is clever enough to figure out that it is large and has a high average number of requests by e.g. seeing that many people use their distinct Google Accounts.
I think what happens is that when they detect too many requests from an ip address in a short interval amount of time they throw a captcha.
Google is in all likelihood using V8 directly and is explicit about who it is. In fact, most legitimate search engine spiders are.
On a serious note - the article mentions repeatedly that many of the techniques can be trivially defeated. The main purpose of the article was to highlight some interesting, unique properties of PhantomJS that most people are not aware of (e.g. stack trace analysis).
Especially considering nowadays you can automate IE with PowerShell [1] . I've used it myself to automate account creation on a legacy system at work.
[1] http://www.youdidwhatwithtsql.com/automating-internet-explor...
There are a number of ways to try to detect and trick them out. I've done some unusual things to attempt to detect them, finding people coming thru proxies as well.
Sure, these simple methods can be defeated by an adversary, but at least judging from my experience, even simple methods you're going to catch some of them. Some hackers ain't that smart. A little surpised Cloudflare for instance wouldn't offer this (PhantomJS blocking), if they don't already. I'm all for hacking, but a site owner you should have the right to block bots if you so choose.