Detecting Chrome headless
antoinevastel.github.io
antoinevastel.github.io
But someone who really wants to do web scraping or anything similar will use a real browser like Firefox or Chrome run it through xvfb and control it using webdriver and maybe expose it through an API. I find these to be almost undetectable.. The only way you can mitigate this is to do more interesting mitigation techniques. Liie IP detection, Captchas, etc.
edit: when I say real browser, I mean running the full browser process including extensions etc.
Practically speaking I don't think IP detection is useful at all these days, and the only Captcha that can't be bypassed is Google's most recent version. The much more successful anti-scraping tactic is to use a sophisticated reverse proxy that analyzes all requests to identify patterns and unusual behavior, indiscriminate from the IP source (because any large scale scraping will come from many IPs anyway).
Having said that, just let them scrape?
Also TOR has public lists of exit nodes (https://torstatus.blutmagie.de/), and unless you exit via a non-listed one you're trivially identifiable.
Lastly, TOR is rather slow, which prohibits using for large-scale scraping tasks. Plus you get switched around different exits and countries, which might break your scraping logic.
There's no such thing as an unlisted exit node, only unlisted entry nodes (bridges). I agree that Tor isn't a good choice for scraping.
If you do some googling, you will see many reasons not to use that extension.
[1]http://www.theverge.com/2015/5/29/8685251/hola-vpn-botnet-se...
Depends on who you are. Hotel and airline websites, for example, are pummeled with scrapers wanting pricing and availability info. Letting them scrape with no limits is costly.
Yes, and they make it very cumbersome to discourage it. I've successfully written scrapers for airlines. It's significantly more difficult than crawling other websites for a few reasons:
1. Session management is wacky - they really like to manage state entirely through cookies, and you typically need to visit a specific set of pages in a specific sequence before you can access the resources you want, like the number of seats available or their prices.
2. Sessions have time limits because anyone who looks at the seats initiates a "soft" reservation on them (this works in a similar way for theatre, concert and movie seating).
3. You don't usually have nice JSON endpoints, so you'll be doing a lot of HTML parsing (which, given the type of HTML you encounter, can be hell).
That is changing. Most scrapers haven't caught on, but the more modern things airlines are pushing out (their mobile sites and native mobile apps) often have really nice REST/JSON api interfaces. The scrapers are often still scraping the old desktop site which will be the last to get that underpinning.
I get that at one point not having this info out there was a good thing for them. But now the car is out of the bag. You can't keep pretending that booking sites don't exist.
There are several airlines where 50+% of all sales are on their own website. Lowest distribution cost.
Also, if you don't have the rock-bottom cheapest prices, you may not want to cater to aggregators that promote those.
All other things being equal, I will strongly prefer buying directly from the airline, because I've personally experienced the shifting of responsibility if things go wrong during a flight.
I've also been denied boarding (not in my home country) because the airline claimed that the OTA (Netflights) didn't pay for my ticket, leaving me to spend the night in the airport and having to book a one-way flight with a different airline the next morning.
That was a horrible experience I'm determined to not repeat.
Bypassing captcha automatically or with captcha bypass services where you pay per captcha completion?
Having said that, it has flaws compared to headless Chrome (more moving parts, terrible security, extra memory usage / dependencies for the JVM to run Selenium) so unless that camouflage is proving crucial, I'd avoid it where possible. I've recently migrated a project to headless Chrome and definitely prefer it.
- Neither Selenium nor Webdriver (at least the Python client, but I assume others too) support HTTPS at all.
- By default it opens browsers configured to accept any SSL certificate. I can see why that's useful for local testing, but it's a terrible choice to default to.
- It logs way too much. Logs every keypress sent to it, including passwords.
At least the latter two are fixable, but not trivially so, and better defaults with easy client-side options to opt into insecure mode would have been much better. The former is not fixable without patching both the client and server and in 2017 it's pretty poor to not even have an option for.And, really, the first two points are clearly contradictory. I'm guessing I just misunderstand what you mean?
Do you mean the communication between the driver? I'm curious why ssl would be important there? Should just be done locally and a standard ssh tunnel can help with any remote encryption you might want.
As for the third point.. That's why we have DMZ's..
On further investigation, I think I was wrong on the second though; it might be that it's built into Chromedriver, not Selenium, which would explain why you didn't have the same issue with Firefox.
SSL between the client & Selenium server would be of benefit to keep any bad actor on the network from man-in-the-middle attacks on it. I'm not familiar with how I'd set up an ssh tunnel for that, I'm happy to believe it can be done, but it'd be a lot easier if it just supported HTTPS to begin with.
HTTPS would be an odd choice, if only because I don't want to do any cert management for my selenium test runner. An encrypted channel makes some sense, but I'm not sure what the best mechanism for the shared secret would be. Ssh kind of gets you there, but is not as straight forward, as you ntoed.
Instead, I'd urge to keep the communication at localhost (or tunneled over SSH, at worst). Preferably with a locked down security model on the network so that you will see any traffic going off.
That can also be easily bypassed (I’ve done that, without actually planning to do so) if you can make your behaviour seem completely human.
So if you have a bot that can browse in a way that seems statistically human you can actually get around that. You can still scrape the same datasets – you just need to keep instances of bots, with all cookies around, and have them browse in certain ways. Classify what category a site might likely belong to, put them into buckets of queues, and have bots pull from queues they’re likely to browse, or search for certain terms they’re likely to search for.
In my case, this all happened by accidents – I had IRC bots for a few channels, each kept a perpetual session and would visit every site that was linked (to get the title) and would be able to search Google by scraping. One day I was accessing a bot remotely, and told it to access a page that was NoCaptcha protected (because the site wasn’t working on my home system), yet it passed the captchas perfectly fine. Tried a few more times, always worked. So I tried figuring out why it worked.
I thought that a client which doesnt run javascript would be a huge red flag.
Ultimately, no protection is unbreakable, there's a work around for almost everything. If your site has thousands of pages (e.g. big online stores, that are common target for spidering) it's probably the best approach to make things as slow and complicated for spider author as possible. That's exactly the same logic like with captchas, they can be broken fairly easy nowadays, but they'll still slow down the spidering rate and pump up the cost.
* machine-generated HTML and CSS identifiers
* randomly inserted, unused HTML elements
* images instead of textMachine generated HTML is a pain, but I wouldn't call that "very difficult" - it's more like "annoying" in the same way that having to parse HTML instead of finding a neat JSON endpoint is.
And I'm not sure how randomly inserted HTML elements would help - if you're already parsing the HTML, you can extract the relevant data. Unless you're trying to parse the HTML with regex, in which case: https://stackoverflow.com/questions/1732348/regex-match-open...
IMHO the best way to stop spiders is to control access to your pages. Require users to register and then login each time to see the data, and then have some pro-active monitoring & counter measures in place, monitoring the patterns per users, and per IPs, and across the whole system for an unexpected increase in activity or bot-like behavior (moving too fast or at unusually uniform speed, following links too sequentially, etc.).
If it's not an option, the next best approach is to make it hard for spider to fetch the whole dataset. Don't let user just browse all the pages, instead force them to use search, it makes it much harder to cover everything (and it will not affect normal users too much). You can, for example, return just the 100 products at once, and ask user to refine the search if there's more products than that. Then you can create some monitoring system to watch over unusual search queries that look like dictionary attacks. I've actually built a system like this for one client and it worked fairly well (combined with a few more tricks they were already using).
Of course, all of this can be tricked too, it's all a game of cat & mouse, trying constantly to outsmart the other side.
Firefox + Chrome both lower the priority of background tabs, and may be doing other tricks so the background tabs can stick around and be switched to quickly.
This is not true. I've seen it done in AWS for less than $2,000 monthly, and I can do it from my home for less than $500 per month (minus the costs of my workstation and networking gear, which you'd amortize over the expected lifetime of the project). I have a home server with 128GB of RAM and an i7-6900K, with 125 static IPs and a bunch of Ubiquiti networking gear. You don't even need that much memory or compute power, but I also use my workstation for other research projects. I use my own static IPs to parallelize without having to sacrifice latency or cede control to a shady proxy farm.
It's a pretty straightforward setup - each static IP is given its own route across the switches from the gateway, and the switches have link aggregated connections for 40GbE bandwidth. You have pub sub and queuing, and each scraping target is its own headless Chrome process. The requests to the target from each process are round robin sent across the available interfaces. Then you've got parsing and a local database.
I'm not a particularly invasive scraper (I make my User Agent deliberately obvious with an explicit way to opt out), but it's really not true that it's prohibitively expensive. This is a pretty cheap setup; I don't even profit from the work, I use it for personal research. If someone was actively selling valuable data, this would absolutely be worth it.
With 128GB of RAM I can only load small amounts into memory for targeted analysis. Processing the rest of the data requires loading it directly from storage. To improve I/O performance I parallelize the transfer across link aggregated ethernet interfaces. Naturally that would cause disk reads to become a bottleneck; to take proper advantage of the network transfer speeds I hold all data in RAID 0 with 7200 RPM drives.
To avoid a block, you need a list of good socks5 proxies
They're not.
> To avoid a block, you need a list of good socks5 proxies
No; furthermore, that leaks your data to a third party and introduces significant latency.
Typically display none will do but you can also have white on white text with no tabindex or off screen absolute positioning, etc.
Wouldn't block specialised scrapers as they would know to avoid that URL/link (though they'd probably work it out anyway), but would still limit more broad crawlers.
I have found that advanced bot mitigation is the single most significant determining factor of PPC ROI, which is a sad commentary on the current state of the paid advertising ecosystem.
To answer your question, based on my personal experience, I'd say on average, for display campaigns (specifically not referring to Google search ads, which tend to have a lower percentage of bot traffic - while Bing is the Wild Wild West)...I'd say overall it's somewhere in the 30% neighborhood. Not all of those have malicious intent, but in PPC, every non-human click on your ads is malicious.
Regardless of niche, realtime bot mitigation is probably the best competitive advantage that one can have in the PPC arena.
I say this as someone who has had to combat this specific technique - I'd suggest that if you believe it works, it's probably because you saw obvious scraping activity stop when you did it, but you were never aware of the more professional scraping that adapted to it or was never caught by it in the first place.
Whenever I've scraped a website, I am extremely careful not to request more pages than I need. A better method for blocking them is to flag requests for resources that do not proceed in a logical manner. For example, if you have an API endpoint that displays the information scrapers want, that endpoint should have a specific "route" through the user interface. If you find requests directly to that resource without first proceeding through the typical UI flow, that is more accurate for identifying a scraper.
This is still not foolproof, because the scraper can just script requests to the requires series of pages in order. But it's a good start for getting rid of most scrapers. The most effective method for getting rid of scrapers is IP agnostic behavior analysis, because it can catch e.g. scrapers trying to parallelize requests that increment across a proxy farm or requests that don't obey typical behavior constraints in the UI.
Because I guess you could defeat webscrapers by making the HTML structurally totally different from what is visually present in the browser window.
It has been a standard tactic to run browsers in purpose built VMs. The last bastion were smarter, more disruptive captchas and rendering performance profiling, but even that got grinded by now.
The situation is even more grave with mobile ads as you have zero opportunity to shove captchas, or rely on single unique per-system ID or cookie that was vetted outside of embedded webkit environment.
http://web.archive.org/web/20121127001253/http://simile.mit....
Demonstrations have usually been by security researchers, but these tiny boards can be used wherever one wants to avoid the labor of repetitive navigating, typing and mouse clicking.
I haven't done it but I think mouse movements could be used too. Add onMouseMove event listeners all over the page and see how they are triggered.
A lot of people access the web through smartphones so only allowing people with mouse movements seems a bit overkill.
http://web.archive.org/web/20121127001253/http://simile.mit....
https://bugzilla.mozilla.org/show_bug.cgi?id=1169290
There's a webdriver protocol mailing list with a long thread that describes how it should be implemented and how people could avoid it; ex. recompile without the feature. I just can't find it right this minute.
I really don't think scraping should fall onto that list.
There isn't even a consensus in the IT world whether or not scraping should be able to be legally restricted.
The author has certainly made his position clear, and I disagree with him too.
There are laws to allow for indexing of contents.
That last one is an interesting one. I think one of the most effective way to deter a scraper might be to just provide an API!
Now if you were using the scraped data to republish (copyright infringement) or use it to gain a competitive advantage (re-pricing in eCommerce comes to mind) that is a different story.
(Though I guess the real-life solution would be both simple and depressing: make an exception for googlebot and don't care about anyone else)
All robots forbidden, except googlebot
After which people keep claiming google search is better, not just given special treatment.
what's re-pricing?
Just let the web be the web, and stop trying to control it.
There are also the scrapers blindly looking for vulnerabilities or other unsavory tactics.
You're just not always in a place to scale to the abuse or build something more complex than some simple heuristic filters.
Often? Based on what data?
I find it much more likely you only often notice aggressive scrapers. That however tells you nothing about the behavior of the average web scraper or web scrapers in general.
Not sure about stock prices (I think it's pretty common to pay for real time data there?).
But I can certainly see sites that have a lot of data for their users facing major bandwidth costs if a lot of people were scraping their data. This type of detection isn't really an answer for that, though, as it's easy to mitigate for a scrapper.
Blocking scraping is like DRM. Don't do it. Use a legal mechanism to deal with copyright infringement, and use acceptable usage policy to deal with heavy users that are using more than their "fair" share of bandwidth.
Still, I agree that if people are going to try detecting headless Chrome, Chrome should strive to thwart that. The attacks in the OP seem like low-hanging fruit; I was expecting something more akin to timing attacks on how long Ready events take to fire given delays from actual rendering. Writing the code to imitate that would be a fun week.
In some cases they don't even want it to behave exactly like the regular browser. As soon as your website uses any client-side state (cookies, IndexedDB, HTTP caching, service workers, local storage) you want to have to an easy "give me a clean and isolated browsing session" switch like Headless offers.
People scraping the web are not the target audience of this.
Mind you Google is extremely interested in bots...
I wouldn't be that surprised if they added webgl support later as well.
Headless Chrome is awesome and such a step up from previous automation tools.
The Chromeless project provides a nice abstraction and received 8k start in its first two weeks on Github: https://github.com/graphcool/chromeless
I guess I disagree with the premise of this article.
How is web scraping fundamental malicious?
What rights/expectations can you have that a publicly accessible website you create must be used by humans only?
I'd say this would be one of the most controversial things to refer to as 'malicious'.
[0] https://arstechnica.com/tech-policy/2017/07/linkedin-its-ill...
As others have mentioned, there is nothing (that I know of) that can thwart a motivated and resourceful scraper.
But in other cases you simply don't want anyone to be able to extract your info automatically. A good example would be e-commerce sites which don't want anyone to be able to scrape their pricing information in large scale and real time.
There are plenty of websites that want scrapers to disappear.
And by releasing headless chrome, they killed off some of the competition. (https://groups.google.com/forum/#!topic/phantomjs/9aI5d-LDuN...)
Of course, Google announces itself as GoogleBot. It wouldn't surprise me if they did a second stealthy crawl to detect cloaking. (But I think they are honest when they say they don't, and just throw cheap human labor at it instead by having people browse suspect sites.)
They actually run secondary tests that aren’t GoogleBot. They’re quite easy to detect on very low traffic sites. If you only have a few hundred users, all of which you know personally, and suddenly over the range of a few hours a few users using Chrome visit the page, while it’s not linked or findable in any search engine, just shortly after users using the googlebot UA visited it, and with certain usage patterns – it’s quite obvious.
Detecting the Android Bouncer’s VM is equally easy, although that only happened by accident because, due to an automated action, my app crashed in that and submitted a crash report that was unusual, and I managed to extract parameters that’d allow detecting it (similar with other android virus scanners), but I only cared about that to be able to split those "devices" into a separate category in the crash tracker (all my apps are GPL licensed, and don’t do anything evil anyway)
Anyway, on a purely technical level, scraping of publicly available content isn't inherently bad unless you're asked to stop, or are scraping so quickly as to cause a service disruption by tying up the target systems. There is nothing malicious about generating normal traffic at the rate of a regular user. The animosity arises from what you plan to do with the data, and whether the entity you're scraping agrees with your usage.
Regardless as others are saying, using complete Chrome or Firefox with webdriver solves all these, right? Is there a way to detect the webdriver extension? That's the only difference I think from a normal browser.
I personally already do a lot of Chrome and Firefox in real desktop environment. I love doing it this way. I know I can mimic a real user. It gives me great comfort.
Still, it's only cheaper if you're doing all the tracking down yourself. And I never meant for that to be my point.
All of them. As soon as you can run some JS code before the page does, every single difference can be monkey-patched. There's no way to distinguish native APIs from fake APIs made by someone that knows all ways of detecting them.
You can just use document.body.
I also suggest to use a data URL instead. E.g. "data:," is an empty plain text file, which, as you can imagine, won't be interpreted as a valid image.
let image = new Image();
image.onerror = () => {
console.log(image.width); // 0 -> headless
};
document.body.appendChild(image);
image.src = 'data:,';
> In case of a vanilla Chrome, the image has a width and height that depends on the zoom of the browserThe zoom doesn't affect this. It's always in CSS "pixels".
at what point does it make more sense for companies to just start offering open APIs or data exports? Obviously it would never make sense for a company who's value IS their data, but for retail platforms, auction sites, forum platforms, etc... that have a scraper problem, it seems like just providing their useful data through a more controlled, and optimized, avenue could be worth it.
The answer is probably "never", it's just something that comes to mind sometimes.
So thinking about how to ward off bots that do go the extra mile makes sense. (From a scrape-protection POV at least)
Cheating an advertiser I'll grant you, but the other two are 100% legitimate.
Since when web scraping considered malicious? Companies like Google are doing billions because they use web scraping.
Google’s current captcha system does track a few of these, but it mostly takes your browsing history, and, if that seems normal, will accept you.
I’ve run a few IRC bots that allowed people to submit Google searches, and would return the first resulting link. They also fetch any link mentioned in IRC channels, execute the JS, and after a timeout of 400ms respond with the current page title.
Both combined – a normal search history, reading a few hundred pages and videos a day per user – apparently are enough that they seem "human", and can pass NoCaptcha.