Detecting Chrome headless, the game goes on
antoinevastel.com
antoinevastel.com
In headless Chrome, the "Accept-Language" header is not sent. In Puppeteer, one can force the header to be sent by doing:
page.setExtraHTTPHeaders({ 'Accept-Language': 'en-US,en;q=0.9' })
However, Puppeteer sends that header as lowercase: accept-language: en-US,en;q=0.9
So it seems the detection is as simply: if there is no 'Accept-Language' header (case-sensitive), then "Headless Chrome"; else, "Not Headless Chrome".This is a completely server-side check, which is why he can say the fpcollect client-side javascript library isn't involved.
Here are some curl commands that demonstrate:
Detected: not headless
curl 'https://arh.antoinevastel.com/bots/areyouheadless' \
-H 'User-Agent: Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/76.0.3803.0 Safari/537.36' \
-H 'Accept-Language: en-US,en;q=0.9'
Detected: headless curl 'https://arh.antoinevastel.com/bots/areyouheadless' \
-H 'User-Agent: Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/76.0.3803.0 Safari/537.36'
Detected: headless curl 'https://arh.antoinevastel.com/bots/areyouheadless' \
-H 'User-Agent: Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/76.0.3803.0 Safari/537.36' \
-H 'accept-language: en-US,en;q=0.9' --lang=en-US,en;q=0.9
You can prove this with the following Puppeteer script: (async () => {
const puppeteer = require('puppeteer');
const browserOpts = {
headless: true,
args: [
'--no-sandbox',
'--disable-setuid-sandbox',
'--user-agent=Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/76.0.3803.0 Safari/537.36',
// THIS IS THE KEY BIT!
'--lang=en-US,en;q=0.9',
],
};
const browser = await puppeteer.launch(browserOpts);
const page = await browser.newPage();
await page.goto('https://arh.antoinevastel.com/bots/areyouheadless');
await page.screenshot({ path: 'areyouheadless.png' });
await browser.close();
})();Recently, we have seen a spike in sites that detect, and block our crawlers with some sort of Javascript we cannot identify. We use headless Chrome and selenium to build out most of our integrations, I'm starting to wonder if the science of blocking scraping is getting more popular...
I don't think what I'm doing is subversive at all, we're running background checks on people, and we can reduce business costs by eliminating error-prone researchers with smart scrapers that run all day.
I don't want to seem like the bad guy here, but what if I wanted to do the opposite of this research? Where do I start? Study the chromium source? Can anyone recommend a few papers?
I'm curious why you'd jump straight to browser detection as the most likely culprit. When I was doing scraping, the far more common case was bot detection by origin and access patterns. It's just very difficult to make an automated scraper look like a residential or business user.
Where do you run your scraping operation? Is it in AWS or some other hosting provider, because that will get you blocked quickly by a lot of sites? Do you rate limit, including adding random jitter to mimic the way a human might use a browser?
There's scraping services available that essentially use a network of browsers on residential connections with their extension installed to get around scraping detection. It's much slower, but it's much more reliable. We also had some success by signing up with a bunch of the VPN providers (PIA, NordVPN, ExpressVPN, etc) and cycling through their servers frequently. Anything to avoid creating patterns that look automated or being tied to an IP that can be blacklisted. I'd start there before I'd worry about hacky javascript detection like in this story being what's tripping you up.
We do have human emulation routines that helped avoid most detection, and that library is decoupled in such a way that we can edit behavior down to the individual site.
Some sites are just so damn good and detecting us and I just don't get it.
The countermeasure would be to have a bunch of humans use the websites in any way they want, totally undirected, then use the totality of that browsing to facilitate your scraping probabilistically. It would be less efficient, but very difficult to catch.
Given that your run this division there is a good chance you are personally liable.
We've had this division for many many years, and before my time we paid another company to do this. There's no legal issues.
Computer security laws are very broad. It doesn't matter if it's just a website that the public can access. If you're accessing it in a matter that they don't want AND you're aware of that, then I struggle to see how your lawyers can justify it.
> Computer hacking is broadly defined as intentionally accesses a computer without authorization or exceeds authorized access.
https://definitions.uslegal.com/c/computer-hacking/
Hiding your user agent because you know they don't want automated retrieval of information is "without authorisation".
Don't think connecting a computer to a private network to suck up subscriber data is comparable to scraping publicly accessible internet content.
First, the accuser needs to, at least, send a cease and desist letter to the accused asking them to stop accessing the protected computer. Second, the accused needs to ignore that request and keep accessing the protected computer.
Is it possible to build a solid CFAA case when those two things do not happen? I cannot find any examples.
https://iapp.org/news/a/can-a-cease-and-desist-notice-create...
Even if you aren't trying to disguise anything, adding some randomness helps avoid one particular bad pattern with operations on a network. I recall the pattern being called "network synchronization" but I can't get good search results for that.
Are you saving money at the expense of the site operator by scraping their site for public records, or are you saving money as well as the site operator?
If you're costing them money to reduce your own bottom line without their express written consent, that makes you "the bad guy". Offsetting costs onto an unwitting, non-consenting third party is an unethical approach to doing business.
I interpret your request as a similar problem to "help me with my homework problem". I could dig up papers and studies, but at the end of the day, you need to go do your homework. Reach out to each municipality and figure out a business arrangement with them that satisfies your needs. It's possible they do not wish you to perform this activity, in which case you will either need to violate their intent for your own profit using scraping or accede to their wishes and stop scraping their municipality. That's your homework as a for-profit business.
We measure the value in FTEs, and when a researcher quits, we do not replace them if the appropriate FTEs have been reached with projects.
It's a major benefit to the business not only because we don't have to pay another employee, but we can reduce training costs, and costs incurred by mistakes. We can also adjust execution of one of these agents, which normally would require rearrangement of work instructions, and retraining.
These are public records, 90% of them do not have integrations for automated systems, and those that do, we utilize. They are typically search boxes with results. We are not circumventing any type of cost that would otherwise be incurred.
We do not log any of the results, store them locally, or maintain any of the PII with each search. If a case was searched 20 minutes ago, and comes up again, we rerun the entire thing just as a human would.
Finally, to your point about 'help me with my homework', I consider posting on the HN forums homework for this type of research. There are a diverse set of talented developers on here with esoteric experience. The fact that an article related to the work I do came up on here, I thought, was an excellent opportunity to seek advice and perspective.
If it reduces market for some consultants, well, sucks to be them, they'd better find a different way of providing value. Not every value needs to be captured and priced. A world in which all value was captured and priced would really suck.
This is the same attitude that says, "why would someone just give away Open Source software when they could build a SaaS business instead?"
Sometimes the answer tells you much more about what skills you need to be hiring. Sometimes they give you a lead.
The fact that some government organizations make it hard to retrieve public records is a flaw in the system. I'd be in favor of a national law requiring all public records to be published in machine-readable form.
In the mean time, it is our civic responsibility to conspire to circumvent these misbehaving public services.
No such funding exists, and municipalities are regularly denied tax increases by their voters for any reason — much less public records publication that would often embarrass and humiliate those same voters.
So in essence you're asking them to cut public services and staffing in order to give hundreds of dollars of IT costs a month to for-profit businesses who can't be bothered to pay some small fraction of their revenue for the costs of delivering those records.
It is our civic responsibility to republish those records for free as citizens. Doing so for profit at the expense of citizens is unethical.
If OP republishes all records received in a freely-downloadable, unrestricted form, then I would happily help them fix their scrapers. They, of course, do not.
The public records are public. Charging for them is, by the above arguments, immoral. Therefore, not only the municipalities but also the businesses profiting from those public records owe us their scraped data, for free, without regard for profit concerns.
Not one for-profit business does so. Why is their immoral action acceptable, when the same action by a municipality is not?
The problem here is that instead of building APIs (or just posting to FTP sites), governments are building offices and funding staff to answer snail mail requests. Or building sophisticated web forms and search engines.
It's obvious how we got to this point (before the internet, you obtained public records by walking into an office) but it's long past time to change. We don't need fancy web forms to search and find data; cut all that out and just provide data in machine readable form to anyone who wants it.
Someone will build a pretty commercial interface to public records data. Chances are, they can do it for less than the 8-figure sum required for UI development in the public sector. Win-win.
Discussing salaries is a taboo created by industries to stifle wages.
https://www.monster.com/career-advice/article/truth-about-di...
Currently? Not off the top of my head. But there was one that scraped municipal records in a large midwest city and made them public for free because they were confusing to get to otherwise.
Unfortunately, the company was bought by a larger company and that portion of what they did was shut down.
If a real-world demand for, say, some GIS data is hundreds of requests per day, then a crawler that comes in with hundreds requests PER MINUTE will obviously stress the infrastructure. Adjusting infrastructure to cope is not an instant process, nor is it a sure thing to begin with given all the budgeting formalities. So your "civic duty" will ultimately result in destruction of these services, because they simply don't have the means to deal with such thoughtless activism.
If the municipality wants to get the information out, this could be a win-win, just like search engines were. Do check for robots.text, though!
Voluntary honor systems don’t work, because there’s no way to compel non-compliers to stop other than standard “anti-attacker arms race” approaches, such as the obstacle described at the head of this thread.
Are you really arguing that the internet would be _more_ accessible if search engines had to reach out to every site they wanted to crawl?
How many companies out there complain about being scraped by Google? How many companies benefit from search-driven traffic?
Naturally, Google didn't want that.
The best example is a large number of unimportant sites that send 429 errors for /robots.txt if they think it's a scraper. A 4xx result for robots.txt is considered to mean no robots.txt for most crawlers. So the website is getting the reverse of what it thought it was getting.
However, from what I've noticed of search results over time, these background check (AND identity verification) sites crawl each other and create a kind of feedback loop, as I've been noticing that some of these pages will falsely report parts of her sibling's background among her own, and falsely flag her as having certain ugly events in her past that don't actually belong to her. This is concerning, as her career area cares a lot about employees having a clean background, and employers using these cheap automated options see cheap, inaccurate results. She has a squeaky clean background with a high credit score and impressive educational credentials, while her sibling has had run-ins with the law and bad debts. I'm concerned about how this will affect her future career prospects.
Beyond background checks, identity verification is a big concern as well. You may have noticed some services ask you to confirm certain facts about your past (street names of where you've lived, schools you attended, jobs and cars you've held). When pulling her credit bureau reports, some of these verifications required confirming facts about her sibling rather than her own in order to gain access.
Like I said, these issues have been fixed with all the "official" record-keeping organizations; however, since the fix, I've been noticing increasing issues with the original mistakes propagating to 3rd-party background-check organizations.
These services cause more problems than they solve, and should require consent, oversight, and civil or criminal penalties associated with a failure to meet high quality standards.
Existing law does not proscribe recklessly sharing damaging false information about people?
Also of note, "Malice would also exist if the acts were done with reckless indifference or deliberate blindness" - https://en.wikipedia.org/wiki/Malice_(law)
First, the process of automating a source is not as simple as 'grab data, send to person that creates the case'.
We have many many layers of precaution and validation both by humans and other automated systems, that helps guarantee accuracy.
On top of this, even public records has reporting rules in the industry. There are dates, specific charges, charge types (Misdemeanor/Felony), and a battery of other rules that the information is processed through in order to ensure we do not report information that we are not allowed.
We always lean to the side of throwing a case to a human. In the circumstance that anything new, unrecognized, or even slightly off happens, we toss the case to a team that processes the information by hand. At that point, we are simply a scraper for information and we cut out the step of having a human order and retrieve results.
We do not go back 40 years. Industry standards dictate that most records older than 7 years are expunged from Employment background checks. And most of our clients don't care about more than 3 years worth, with exceptions like Murder, Federal Crimes, and some obviously heinous things.
We also run a number of other tests, outside of public records to provide full background data. We have integrations with major labs to schedule drug screens, we allow those who are having a background check run on them to fill out an application to provide reasoning and information from their point of view to allow customers to empathize with an employee.
We also have a robust dispute system. The person having a background check run on them receives the report before the client requesting it in order to review the results and dispute anything they find wrong. These cases are always handled by a human, and often involve intensive research, no cost spared, to ensure the accuracy of the report.
There's a plethora of other things I'm missing, but if you have any specific questions, I'm happy to answer.
*EDIT
To clarify, there is a lot of information in public records. It isn't unclear or ambiguous at all. Motor Vehicle and Court records are extremely in-depth and spare no detail.
As a private person, we only have access to court documents on a state or county base. Any central database we have access to would be made my scrapers.
The right to be forgotten is alive and well most of the time, 90% of our clients don't observe information further back than a few years. I feel like that is a fair assessment of someone's behavior.
There is a point where data collection becomes unethical, and making everything fine as long as it isn't legal makes for a shitty society. (i.e. legislating behavior should be a last resort not a first judgement on right and wrong)
I don't know precisely where that point is, but automated scraping of social media probably is past (automated scraping of judicial records? probably ok)
I think the extent and reason for the checks aren't apparent. So I'll give a few examples where we have high volume and I hope that will enlighten you as to the reason why there are so many players in the industry.
The highest volume checks are around the medical and teaching fields. We often run 6-month, to one year recurring checks on teachers and doctors to ensure licenses and certifications are still active. As well as necessary immunizations to work in their environments.
Do you expect a low margin industry like teaching to staff a full time employee to do nothing but run background checks? They want them done and the schools have access to the information, it's just much easier for them to pay us a few dollars an employee and get a nice report than do the legwork themselves.
Additionally, incurring the cost of access for the relevant data is a barrier for companies without a bunch of cash laying around.
We don't solicit companies with incriminating information about their employees, it's a necessary part to a safe environment.
What isn't harmless is gathering information about the private lives of people (even when done in the public eye) in ways that are difficult, labor intensive, or impossible without automation.
A previous company I worked for aggregated publicly recorded mortgage data. The mortgage data was scraped from municipal sites on a nightly basis because it was not available as a bulk download or purchasable option.
We had requested on several occasions for a service we could pay for in order to get a bulk download of this data, but the municipalities did not have the know how to provide this as were using systems from a private vendor that were prohibitively expensive for them to request modifications. As a result, we worked hand in glove with the municipalities to ensure we were not stressing their infrastructure when we did this scraping, and I think that's the best we were able to do in this case.
If you do want to stick with Selenium, you're better studying the chromedriver source than Chromium itself.
Not sure if people do this sort of things nowadays.
i do my scraping just for myself. maybe if i would scale it up they would detect me.
redPill: function() {
for (var e = performance.now(), n = 0, t = 0, r = [], o = performance.now(); o - e < 50; o = performance.now()) r.push(Math.floor(1e6 * Math.random())), r.pop(), n++;
e = performance.now();
for (var a = performance.now(); a - e < 50; a = performance.now()) localStorage.setItem("0", "constant string"), localStorage.removeItem("0"), t++;
return 1e3 * Math.round(t / n)
},
redPill2: function() {
function e(n, t) {
return n < 1e-8 ? t : n < t ? e(t - Math.floor(t / n) * n, n) : n == t ? n : e(t, n)
}
for (var n = performance.now() / 1e3, t = performance.now() / 1e3 - n, r = 0; r < 10; r++) t = e(t, performance.now() / 1e3 - n);
return Math.round(1 / t)
},
redPill3: function() {
var e = void 0;
try {
for (var n = "", t = [Math.abs, Math.acos, Math.asin, Math.atanh, Math.cbrt, Math.exp, Math.random, Math.round, Math.sqrt, isFinite, isNaN, parseFloat, parseInt, JSON.parse], r = 0; r < t.length; r++) {
var o = [],
a = 0,
i = performance.now(),
c = 0,
u = 0;
if (void 0 !== t[r]) {
for (c = 0; c < 1e3 && a < .6; c++) {
for (var d = performance.now(), s = 0; s < 4e3; s++) t[r](3.14);
var m = performance.now();
o.push(Math.round(1e3 * (m - d))), a = m - i
}
var l = o.sort();
u = l[Math.floor(l.length / 2)] / 5
}
n = n + u + ","
}
e = n
} catch (t) {
e = "error"
}
return e
}
};If I load the page source in Chrome, it already includes the "You are not Chrome headless" message, but when I run it in a scraper I maintain, the page source loads with the "You are Chrome headless" message, even without running any Javascript.
Those are averages of multiple runs on a Core i7-8550U running Chromium 75.0.3770.90 on Ubuntu 19.04.
isNan and isFinite are much slower in headless mode, but other functions like parseFloat and parseInt aren't. My guess is that the backend is comparing the relative times that certain functions take. If isNan and isFinite take the same time as parseFloat, then you're not in headless mode. If those functions take 6x longer than parseFloat, you're in headless mode.
I don't know if this holds true for non x86 architectures or other platforms.
Unexpectedly, it turned out that Accept header was perfect for this. The final chart was this:
https://i.imgur.com/ZA8qD8t.png
("link" means clicking on an URL or entering it manually; "embedded" means <img> tag)
Makes me wonder whether Accept header is still useful for fingerprinting in general, and distinguishing between headless and headful(?) browsers in particular.
People that are knowledgeable enough will deep dive into the webpage, but for everyone else, expect disappointment.
It’s sad to have a smart guy like this dedicating his academic career to something this inconsequential. Anyone with enough incentive is going to be able to defeat any technique this guy dreams up. Anyone who might benefit on paper from detecting a headless browser isn’t going to want this because any possibility of a false positive is a missed impression, or a missed sales opportunity, or an ADA lawsuit (US), or an angry customer.
Edit: According to other commenters there are checks in the included version of the library which are not in the release version.
If it were doing something like using CSS being non-blocking (? I don't know that it is) that's a server side detection .. but that would seem to work even against spoofing.
But he says if you spoof another Chrome-based browser (Safari) he can't tell. So he's looking first at UA?? That's weird.
Moreover, I'm not even sure how useful would this community testing be academic-wise. Black-box testing is great stuff for CTF competitions. But any decent academic venue would dismiss systems that can't withstand white-box testing as security-through-obscurity.
The thing is, we would gladly pay these companies for an API or even just a periodic data-dump of what we need. We've even offered to some of them to write and maintain the API for them. They're not interested, for various industry-specific reasons.
I often wonder how much developer time and money are wasted in total between them blocking and devs working around their blocks.
Wanting to give out limited free samples inevitably leads to making sure you are giving out samples to people and not bots and not too much to each person, and that leads to user tracking.
Compare with the arms race between newspapers and incognito mode:
https://www.blog.google/outreach-initiatives/google-news-ini...
=> Using Chrome directly is slower, but undetectable.
Thankfully they're also so technologically slow that they never change the websites or do any kind of headless detection. Its works, and allows us to offer automated [process] to our customers, but it seems so fragile. Just give us a damn API.
He basically says he's inspecting user agents:
> Under the hood, I only verify if browsers pretending to be Chromium-based are who they pretend to be. Thus, if your Chrome headless pretends to be Safari, I won’t catch it with my technique.
Maybe I should apply for a PhD too.
So if you take a Headless Chrome instance but change the User-Agent to match that of a normal Chrome, does the detector think it's not headless?
> It does not use detection any of techniques presented in these blog posts (post 1,post 2) or in the Fp-Scanner library
That's from the linked test page https://arh.antoinevastel.com/bots/areyouheadless
The studies in which I’ve participated always start with a statement of what they are generally looking for in a participant. You then take a survey that confirms if you are qualified. You then given a release to sign (and keep a copy of), which states what you’ll be doing, and providing an IRB contact. You then go through the study.
At the end of your participation, you are asked “What do you think the study is about?”, and then you were told the real purpose of the study. Eventually the paper(s) is/are published, with hypothesis, methodology, and results.
This seems similar: You decide if you want to participate, and are participating; the only thing that’s missing is the final paper.