New headless Chrome has been released and has a near-perfect browser fingerprint
antoinevastel.com
antoinevastel.com
Edit: Please also note that we have not released New Headless yet. We "merely" landed the source code.
Its bizarre to ask a client side program to implement server-side controls for users you want to allow on your site but throttle.
At least I've never gotten detected.
The best detection I've come across so far (i.e. before this release) has just required I run headless Chrome in headed mode. Granted, I don't do a ton of scraping -- mostly just pulling data out of websites so that I can play with it in aggregate using more civilized tools.
[1]: https://github.com/berstend/puppeteer-extra/tree/master/pack...
This is an awful world, one designed to reinforce class divide and protect the entrenched and the rich by deliberately handicapping easily-accessible tools, because of a few bad actors. It creates a world where the code for literally everything is the most hideously complex version of itself because it is riddled with constant checks, phone-homes, and arbitrary usage limits. It further pushes us towards a disempowering future where our computing is limited exclusively to appliance-like devices whos inner workings are controlled for it. It stands against the very principle of general-purpose computing.
And implications of a question aren't either. Just your imagined implications. Questions aren't bad.
While my immediate response was the same as yours, I think this actually won't really change much in the way of bad actors.
It's unfortunate, but basic controls (such as throttling, etc) are pretty much a floor-required feature - one way to avoid this burden is to do things like use 3rd party idp (aka google login). I'm not happy with the state of things but I don't think headless will particularly contribute to a material increase in abuse cases.
In 2000 sites were running where code has been precisely made such way DDoS attack was impossible. Now it is heckin sauce of js malware obfuscated proprietary code.
If your site like this, you deserved it. Cloudflare and such companies just need your money for solving 5-minutes problem like AWF that is just a regex, and you have limits even for user agent filtering, lol.
Stop making shitcode and learn HTTP and TCP/IP theory, and you will make antispam filter that is 200% better than any cloudflare shit that is simply malware that runs cryptominer as a "IUAM" mode for their own benefit and you even pay for it.
Why go through the exercise, one may ask? I believe it would be a critical thinking exercise to improve Headless even more while giving website maintainers a way to opt out of receiving traffic from it. If not your team, have you reached out to see if people from project zero would take on that challenge in their abundance of spare time? [1]
Well, Headless is open source, which means anybody could build a Headless version with such a property set to "I am a human, trust me!" and employ such a modified binary ... ;-)
My own personal method for my silly hobby sites is just to put passwords on things with an auth prompt delay.
Almost all of those are things are outside of the scope of the browser itself. And anyone doing serious bot attacks already have scripts/forks that modify these signals. I don't see how the chrome team could do much to help stop that at that level.
Or if I put my evil corp hat on, the incentive could be that they make puzzles that only Headless can get around and all other bots become trivial to block and obsolete by even the least knowledgeable hobbyist. Perhaps Google release Nginx, Apache HTTPD, Apache Traffic Server, Envoy and HAProxy modules that only Headless can get around and all other bots internet-wide are entirely silenced. Chrome becomes the one and only bot to rule them all.
I suppose that Google going through that exercise would mean that they get market dominance on bot gathering data and anyone not using Chrome Headless would be unable to obtain freebie data. This could enable future features whatever that may be. readjusts hat One future feature could be auto-discovery of Google DNS and Google proxies in GCP so they can learn about new data sources through crowd-sourcing thus making their big-data sets more complete and their machine learning more powerful. Developers could block the proxies or compile them out but as we know most people are too lazy to do this and many won't care.
Another advantage would be that eventually the only bots abusing Google would be bots using their code and they would know how to detect and deal with as they would implement their own open source anti-bot modules in their web servers, load balancers, etc...
There are more obscure ideas but I am doffing the hat before the hat-wraiths sense it.
The push for this is starting with adult content [1] but the goal posts could easily be mounted on train car with a very long and smooth train track that only goes downhill.
You know what? The Internet Is For End Users [1]. If we're going to cite an RFC, it should be RFC 8890. Not having a better headless Chrome would be a violation of the most basic principles of the internet.
There are some cases where automation can get out of hand, but blocking these efforts should not come at user expense. So says the RFC8890, and a general collective belief/hum-in-the-room. The availability of a good browser like Chrome helping should not be an issue, given how many other ways bad players have to go too far & cause harm to sites. The people who have to deal with this are not the priority & this doesn't radically change their troubles; this radically helps end users wishing to exercise agency though.
In most cases being able to script & automate a site is a completely primitive user-agency, of no special regard. Headless Chrome being a somewhat tolerable way of doing that scripting is 100% morale, correct. It greatly assists us in fulfilling a primary & clear overarching purpose of the internet: to be for end users.
I wish I could say I cannot believe the complaining & whinining & snivelling, the pretentious-nonsense/acting-offended that Chrome would dare help make good automation. I wish I could say I don't think this crowd recognizes nor comprehends the basic purpose of the internet, but again, I think I know better; I suspect they do but their protests are disingenous, that they have allied their hearts with darker forces, against the user.
This is flawed reasoning. Just because we can't eliminate abuse from headless browsers that doesn't mean we shouldn't work to reduce it. Finding such a modified binary or making it yourself is additional friction that will cause less of these bots to exist. Some people may not care if a website is able to block them or not or some people may not decided to do the work to read the robots.txt. By implementing these capabilites into the product by default you are making the web ecosystem a better place wit less abuse. You are right that someone could make a version without the antiabuse parts, but surely that fork will be less popular and less used.
>Why should we make a distinction between humans and machines?
Because machines can be used to abuse a site at a scale that humans can't. Site owners want to protect their site against abuse.
I would hope that Google's robots would not be programmed to lie to me, but would be honest.
If robots are required to be honest, then I have a choice to serve them or not. If they are not honest, I do not have a choice.
Should bots have the legal right to say they are human?
For example - if Google Inc is visiting a web page to collect information about it using a headless bowser, and the server asks - are you a bot - should Google be legally or ethically allowed to answer no? (declarations in headers could remove the need for question/answer chatter.)
(I want to pre-empt dismissing this line of questioning via 'what if Google wants to know how the site will be served to a human for better search results because google could include a specific header for that, eg "I am a bot, but request that you serve the version of this page served to humans". It would be up to the server to honor or reject that request.)
The defaults Google choose have compounding effects in our society. If you make it "normal" for bots to pretend to be human, the industry has minimal pressure to hold any standard above what you do, and better norms may never appear, or be delayed by a decade. The alternative is to be thoughtful today to try to create a better world.
Using the word "new" in naming conventions is the most moronic and shortsighted way to name things in something that is quite obviously going to be changing in the somewhat near future.
--headless
--headless=new
--headless=chrome
And each mean something different - but what?
Not documented, very frustrating.
Can you explain the difference between each of the above arguments?
--headless enables the default, which is "chrome".
--headless=new enables "new".
The new Chrome headless certainly purports to be "just Chrome" "without actually rendering." One of the notable differences in the new headless mode is that it at least shows the stock/built-in extensions. From the submission:
> Similarly, when it comes to plugins, the old headless Chrome used to return no plugins with navigator.plugins, which is a technique that used to be exploited for detection when Headless Chrome got released 6 years ago, cf this blog post. The new headless Chrome returns the same plugins as a headful Chrome, and that’s the same for the mimeTypes obtained with navigator.mimeTypes:
Maybe perhaps the new headless is faking it, but my impression is that extensions definitely work as normal in the new headless Chrome. How or whether they worked before is another very very interesting question I'd like answers to.
I do wish the AMA dev had actually replied to this. My hope is that this wasn't an issue before (but default plugins just weren't installed, and now they are, just to alter fingerprinting), and that now the situation is unchanged but default plugins are installed.
[1] https://stackoverflow.com/questions/16800696/how-install-crx...
It looks like the new headless mode does support extensions.
Is there a list, an explanation anywhere?
One thing I'm hoping for (but have heard it would require extensive rejigging of almost absolutely everything) is Extensions support in this new headless.
However, if I'm reading the winds, it seems as if things might be going there, because:
- Tamper scripts now work on Firefox mobile
- Non-webkit iOS browsers are in the works
- It's technically possible to "shim" much of the chrome.extension APIs using RDP (the low-level protocol that pptr and its ilk are based on) which would lead essentially to a "parallel extensions runtime" and "alt-Webstore" with less restrictions, something which Google may not look merrily upon
Anyway, back to "headless detection", for the remote isolated browser, I have been using an extensive bot detection evasion script that proxied many of the normal properties on navigator (like plugins, etc), and tested extensively against detectors like luca.gg/headless^1
Interestingly one of the most effective way to defeat "first wave" / non-sophisticated bots used to be simply throwing up a JS modal (alert, confirm, prompt) -- for the convenient way it kills the JS runtime until dismissed, and how you have to explicitly dismiss it.
It's "Right to read"[2] all over again.
[1] https://www.ietf.org/archive/id/draft-private-access-tokens-...
There are other kinds of non-invasive bot management you can do as well, however, due to various reasons I'm not in a position to talk about it. A few other methods are mentioned at the end of the post being discussed[2].
[1] https://en.wikipedia.org/wiki/Hashcash
[2] https://antoinevastel.com/bot%20detection/2023/02/19/new-hea...
It was done super fast.. one can't help but think that Google pull all the levers they had at Apple/Mozilla to made sure the first viable alternative to advertisement was killed before it was born. But I think as a side effect it make PoW might be sort of impossible?
I don't really know how to mining "fingerprinting" works exactly - so would be curious to know if I'm wrong
1) It was almost exclusively used for malicious purposes. Very few legitimate web sites used cryptominers, and it was never considered a viable substitute for display advertising; it was primarily deployed on hacked web sites. Browser vendors were relatively slow to react; many of the first movers were actually antivirus/antimalware vendors adding blocks on cryptominer scripts and domains.
2) The most popular cryptominer scripts, like Coinhive, all mined the Monero coin. (Most other cryptocurrencies were impractical to mine without hardware acceleration.) Monero prices were at an all-time high at the time; when Monero prices crashed in late 2018, the revenue from running cryptominer scripts dropped dramatically, making these scripts much less profitable to run. (This is ultimately what led Coinhive to shut down.)
The hacked site boogieman felt overblown (and from what you're saying it sounds like if would have died out anyway). I'm sure it happened, but at least personally I never once came across it. Or if I did, then my CPU spun a bit more and I didn't notice. No real harm done.
More fundamentally we're now in territory where the browser vendors get to decide what javascript is okay to run and which isn't.
Anyway, it's just complaining into the ether :) it is what it is. thanks for the context of the market forces and antivirus companies
Coinhive was live from 2017 - 2019, and it basically ran the whole course from exciting new tech to widely abused to dead over those two years. I don't think it needed more time.
> The hacked site boogieman felt overblown...
Troy Hunt acquired several of the Coinhive domains in 2021 -- two years after the service shut down -- and it was still getting hundreds of thousands of requests a day, mostly from compromised web sites and/or infected routers. It was a serious problem, albeit one which mostly affected smaller and poorly maintained web sites.
https://www.troyhunt.com/i-now-own-the-coinhive-domain-heres...
[1] https://storage.googleapis.com/pub-tools-public-publication-...
Semantic quibble: it's less "proof of work" and more "proof of hardware+work". Or, as they call it, hardware-bound proof of work. The reason you can't offload the challenge to a more powerful device is that they rely on identifying stable differences for each device class that ultimately trace down to the hardware they're running on.
It would be easy to point to the irony of saying "instead of supporting Cloudflare's proposals for PATs, use their CDN product for brute force protection" but on the other hand, they employ a lot of experts in this space and might see the writing on the wall in an increasingly adversarial public internet.
It’s almost as if the conspiracy theory of Cloudflare acting as an arm of the US government and helping in the centralization of the internet is actually true.
i'm honestly asking, not just trying to disprove you. this is a real problem i have right now. ideally i'd get all my thousands of old, never-updated but dynamically generated pages moved over to some static host, but that's work and if i could just put some proxy in front to solve this for me i'd be pretty happy. but afaik, nothing actually solves this.
I've used this trick to survive surprise spikes of traffic in multiple projects for years.
Doesn't help for applications where your backend needs to be involved in serving every request, but WordPress blogs serving static content are a great example of something where that technique DOES work.
It's a blog. Blogs are not complex. Why is your blog's database so awfully designed that 100 pages a second causes it to fall over?
> leading to real, legitimate users being unable to use my site?
You assume that a scraper is not a legitimate user. I argue otherwise. If you don't want a scraper to use your site then put your site behind a paywall.
> What if it’s not a legitimate scraper but someone with hundreds of proxies uses them to DDOS my site for days?
If it's a network bandwidth problem, then a reverse proxy (eg, CDN) solves that.
> Should I sacrifice my uptime to protect the freedom of those unwilling to attest that they’re running on real hardware?
All software runs on real hardware. What is your exact question?
I am accessing this site in a virtual machine. I could be doing it with a headless browser. Why does that matter at all?
Re: content scraping -- I was an indie web dev of a sort for a while and people always ask this question, and the answer is it's impossible to stop. Not even Facebook or big content sites like CNet or The Verge can stop it. At the bottom of it, you can just access the site in a browser and save the source. Content scraping is a rephrasing of "viewing content even just once". Stopping it is antithetical to the web and technologically infeasible.
A fake TPM would be useless for security but just fine for fooling websites that there is a real human at the computer.
> Computer programs can use a TPM to authenticate hardware devices, since each TPM chip has a unique and secret Endorsement Key (EK) burned in as it is produced.
That EK is signed by the TPM manufacturer, and so it’s likely they’ll only trust the keys of physical TPM manufacturers. Good luck forging that in software.
I doubt websites will go to the trouble to keep a list of approved TPMs. It's the SSL root certs nightmare all over again and even worse. No one is going to want to deal with managing a whole new giant list of devices, having fire drill updates to revoke compromised ones, etc.
The screenshotting uses puppeteer and chromium and a read-only session to impersonate the user and screenshot their dashboard.
It uses the old version of chromium and there were many gotchas that required a lot of extra scaffolding to actually render ours and other websites like they would on my laptop. This will hopefully make it easier for us to maintain once implemented.
Either they have a real TPM with a real nvidia graphics card able to decrypt content with a real serial number... Or they don't...
If one graphics card or TPM serial number starts acting bot-like, you can ban just that one.
Some sites won’t care, but for some this will be too high a price for avoiding headless bots.
Sites that use it get my anti-traffic. I don't buy, support, or condone DRM'd media and I actively disable EME on every browser I come across...
There are sites that commercially distribute DRMed video content; say, Netflix. They have a large audience, and they care, whether me and you like it or not.
Am I missing anything?
[1]: https://developer.mozilla.org/en-US/docs/Web/API/Navigator/r...
All other configurations use L3 which is a shared key, e.g. provided by ChromeCDM as it runs entirely on the CPU - which is why Netflix content also works under Linux, albeit L3 is limited to 720p (or 1080p with browser extensions).
Given Chrome's massive browser market share, I'm not sure whether enabling DRM adds anything meaningful to the fingerprint - i.e. I don't think it's possible to revoke an L3 key without pushing out a new version of the CDM to all users of that browser, as has happened once before with Chrome.
FWIW I've tested Widevine L3 decryption works using a ”headless” docker container running Chrome. The only caveat to add is that Chrome must not be started with --headless, but you don't need a real GPU either, Xvfb works just fine.
this is good, but it would also be helpful if you supported the anti DRM movement. Some people have developed ways to get around certain DRM such was Widevine, from dumping your own CDM to Widevine proxy. Just ignoring the problem is not going to make it go away. Over the last two years DRM use for streaming content has increased significantly. If you want to really help, I would look into contributing code to these projects, or donations.
It does nothing to dissuade content gatekeepers from employing restrictive DRM on their sites.
Anti-DRM would be avoiding anything that gives money to those that employ DRM to incentivize the removal of the DRM. Frankly, flat out piracy (streaming ripped content) is more likely to result in the removal of DRM than making it appear that the DRM is working well for the provider.
Blaming humans for desiring privacy is bad. No one here is "trying to behave like a bot".
Exaggerated example: "Oh, you don't want to show me, a random stranger on the internet, your ID? You are behaving like a crook!"
and there seems to be significant overlap between the people who think enabling bots is morally wrong, and people who think fingerprinting is morally wrong. if you value privacy, you have to value privacy for all web users even before you've collected enough data to determine whether that web user is a real person or not.
I don't think you can get the serial number, though?
(And if there was an API for this it wouldn't be a passive one, which makes it inapplicable for fingerprinting)
If so, why isn't this used as an immutable ever-cookie that can't be deleted?
A TPM can attest that some measurements were done with it and it can attest that it comes from vendor X. You can block an entire vendor if they don’t behave but not individual TPMs via remote attestation.
You can use a scheme in which you can set up an „identity“ on first use and then on next use authenticate the same identity. But that identity is kinda per use case.
The article mentions that to use this you need to specify the --headless=new flag.
I know that to set the headless flag i can just use this code:
from selenium.webdriver.chrome.options import Options
options = Options()
options.headless = True
But how would I specify the new part of the flag/option?Basically omit options.headless and use options.add_argument("--headless=new") instead.
But there are still plenty of legitimate use cases for wanting a headless browser that perfectly replicates a normal browser environment. The obvious ones are automated frontend testing tools like https://playwright.dev/
> It’s important to leverage other signals such as:
>
> * Behavior (client-side and server-side)
> * Different kinds of reputations (IP, sessions, user)
> * Proxy detection, in particular, residential proxy detection
> * Contextual information: time of the day, country, etc
> * TLS fingerprinting.
Having a headless browser that behaves exactly like a normal one is tremendously useful for making things. And people who really *need* to block bots also need to contend with "mechanical turk" style attackers anyway. These techniques are also very useful against that approach, which still may be cheaper than making an undetectable bot even with a near-perfect Chrome fingerprint available headless.
Imagine count of false positives.
> * Different kinds of reputations (IP, sessions, user)
Almost all use mobile network right now. One IP can be sticked to thousands of users. Imagine count of false positives.
> * Proxy detection, in particular, residential proxy detection
Most residential proxy are just common ISP ips bought by a face or led by botnet. Imagine false positives of simple home users that are IP ranged like on 4chan.
> * Contextual information: time of the day, country, etc
script.execute(() => navigator.dateOffset = Math.random()...) script.execute(() => navigator.country = Math.random()...) script.execute(() => navigator.etc = Math.random()...)
> * TLS fingerprinting.
Imagine count of false positives, especially because there are 4 common tls fingerprints across browsers.
Just cope and seethe that your antispam filters will never work, antibot measures are fail. Cloudflare turnstile is fail. Bots won as usual.
Right. But it will be massively used just for that.
You might consider looking into some resources on Intent vs Impact (eg, [0]).
IMNSHO, anyone working in tech has a responsibility to consider what their creations can be used for, in addition to what they intend them to be used for. There's just too much potential for scalability of nefarious behavior to do otherwise.
In all seriousness, despite intentions, and I do love headless mode for actual integration tests with Webdriver, it’s no exaggeration to say that it is likely the single greatest avenue for bots and spam enablement across the entire internet, and imo is probably net Bad.
Until now, bots could simply use headful mode to achieve the same effect that is now made available through the new headless implementation.
I don't like the way they pretend(ed) to send funds to websites using their cryptocurrency services, though. Good software, sketchy company.
Or for all intents and purposes 'noise' traffic.
It'd be nice for the powers that be develop an anonymous cookie standard to allow people to flag themselves as 'humans' without enabling the host to know anything about them.
We are fighting wars over problems that we have created for ourselves.
If I ever detect that site uses "anonymous" cookies without my consent, well you will have to pay a lot of $$ for me as a compensation. Enjoy, try your luck, man. I need money, anyways and I love jurisdiction.
Move on.
It fixes some nagging compatibilities with certain websites. I don't bother with anti-bot mitigations, and I don't expect this to be useful in that regard. commercial Anti-Bot doesn't care about how much you spoof your browser fingerprint.
feel free to AMA
fyi, The new version just helps with compatibility, I have not seen it impact anti-bot detection at all.
I configured everything (login & scraping) and start fetching data using your serivce. Then I discoverd that you have to login to load the dashboard but the their get API only requires the serial number to fetch the data .
So, any website on the Internets can know how many plugins my browser has? Ridiculously!
https://developer.mozilla.org/en-US/docs/Web/API/Navigator/p...
If anyone has a working script to login and perform simple task on one of these sites, please share it.
Any guesses?
And maybe, but that will make enduser suffer more (as always), as more false-positives will be caught.
"Dear British Airways, I booked with SAS instead because you assumed a Linux user with Firefox was a bot."
(Or maybe it was the other way round, I forgot.)
...and that's from 3 years ago
https://antoinevastel.com/javascript/2020/02/09/detecting-we...
So far, the play audio option are kind of weird, specially if you're hard of hearing.
There are some less obvious aspects to this that matter a lot in practice:
1. You have to force the code to actually run inside a real browser in the first place, not simply inside a fast emulator that sends back a clean response. This is by itself a big part of the challenge.
2. Doing so is useful even if you miss some automated browsers, because adversaries are often CPU and RAM constrained in ways you may not expect.
3. You have to do something sensible if the User-Agent claims to be something obscure, old or alternatively, too new for you to have seen before.
4. The signals have to be well protected, otherwise bot authors will just read your JS to see what they have to patch next. Signal collection and obfuscation work best when the two are tightly integrated together.
These days there are quite a few companies doing JS based bot detection but I noticed from write-ups by reverse engineers that they don't seem to be obfuscating what they're doing as well as they could. It's like they heard that a custom VM is a good form of obfuscation but missed some of the reasons why. I wrote a bit about why the pattern is actually useful a month ago when TikTok's bot detector was being blogged about:
https://www.reddit.com/r/programming/comments/10755l2/revers...
tl;dr you want to use a mesh oriented obfuscation and a custom VM makes that easier. It's a means, not an end.
Ad: Occasionally I do private consulting on this topic, mostly for tech firms. Bot detectors tend to be either something home-grown by tech/social networking firms, or these days sold as a service by companies like DataDome, HUMAN etc. Companies that want to own their anti-abuse stack have to start from scratch every time, and often end up with something subpar because it's very difficult to hire for this set of skills. You often end up hiring people with a generic ML background but then they struggle to obtain good enough signals and the model produces noise. You do want some ML in the mix (or just statistics) to establish a base level of protection and to ensure that when bots are caught their resources are burned, but it's not enough by itself anymore. I offer training courses on how to construct high quality JS anti-bot systems and am thinking of maybe in future offering a reference codebase you can license and then fork. If anyone reading this is interested, drop me an email: mike@plan99.net
But I guess there are all kind of purposes, some benign some nefarious, and that they somehow influence the bot operation and detection.
It's a whole industry.
> If we consider a user base of ~175 users, and a minimum bot price of 200 euros (175 users x 200 euros), then the bot developers made at least 35K euros (~$37K USD) in initial bot sales.
https://datadome.co/threat-research/inside-sneaker-bot-busin...
This sounds like a solid product / startup idea to me. I worked on spambot detection in a previous job and it's not at all trivial to solve. Though we were specifically interested in detecting the abusive use of bots, not bots in general, so I focused simply on detecting unusual resource consumption rather than fingerprinting.
JS sounds like a bad match for this task. I perform similar checks from the backend with http headers and Python.
Is there a compelling reason to stick with JS despite the added complexity of obfuscation?
Edit: My use case is different than yours as it's part of a pid-free analytics application. However, bot detection is still an important component of that product.
Very true. Capturing, processing, and storing analytics data long-term is expensive. If I eliminate even 50% of that noise, the savings will be worth it.
I'm attempting to identify the bulk of bots with http headers and real-time session monitoring. I also have an unauthorized list (known bad actors) and an ignore list (search bots, etc.). It works pretty well but definitely doesn't begin address the problem as a whole (from a security perspective).
It's an interesting and complex topic.
I like automating menial tasks in shitty web UIs (i.e. clearing out a list of sessions/search history/ad providers that only allow removing a single entry at a time). Simply using Firefox also gets flagged by a lot of these shitty bot detection services. I've never seen them do any useful work.
The only exception is maybe reCAPTCHA or Cloudflare's alternative; that seems to be quite good at catching actual bots, but I do hate most websites that use them because in Firefox you end up clicking on boats twenty times. They're also trivially bypassed by delegating your spamming to click farms, as 1000 minimum wage workers in a faraway country can be cheaper than paying for dev time to work around the minor nuisances of bot detection.
Here comes the sales pitch....