Fighting API bots with Cloudflare's invisible turnstile
troyhunt.com
troyhunt.com
It's a long .gif file which shows that cloudflare's website loads just fine, but HIBP is unusable. Thanks Troy and Cloudflare for making this (free) service unusable. It's free, so I shouldn't expect that it works, anyway. Chrome on Linux and no VPN, fwiw.
> This post may contain erotic or adult imagery.
Yeah it’s fucked alright.
Just out of interest, have you checked if installing Cloudflare's Privacy Pass bypasses the issue?
I’ve just figured it was a fully intentional tarpit, and that I scored too high on some bot heuristic.
I can think of different ways how this could go wrong (i.e. Cloudflare tries to forward the POST request and CSRF/WAF firewalls block the unexpected Cloudflare origin).
Let’s say you run a SaaS using Cloudflare. You may be extremely happy you block 10000 bot requests for every false positive that’s a real human. But let’s say that false positive was a potential customer that would only have paid if they weren’t blocked, and now you just saved less than a penny in server costs to lose hundreds of dollars of lifetime value from that customer.
Sure if you run a free service use Cloudflare. Give in to centralizing the web more, supporting more censorship, and annoying the hell out of a ton of people in the process. But if you’re making money, I don’t see why you wouldn’t have authentication tied to paying users for anything of value, or think of bot traffic as a cost of doing business.
It sucks as a real person using a VPN, but having your website be overwhelmed by bots suck more than a few VPN users having trouble using your website.
If Cloudflare didn't exist, websites would probably just block VPN IPs like streaming services do.
Similarly, say an e-commerce business releases a limited edition product. Many users won't end up getting it anyway so blocking a few users is usually a much better experience than letting bots buy the product for resale later.
On the other hand, it's absolutely infuriating when blogs/search engines come up with these.
The problem is that economics are working on the opposite way. We are not happy blocking 10000 bot requests. We are happy blocking 1 bot that would make millions of requests. This means that it sometimes ok to loose few customers who would pay hundreds of dollars per month each, if it allows us to block one bot who would cost thousands.
Bots have the bad habits of targeting the most expensive parts of the system, if there is one query that is hard/expensive/could be abused, that’s the one that is going to be targeted in priority.
If I meet a "human check", I quickly decide whether it is worth me solving it, or just close the tab. I could imagine 9% of people just giving up. Some of these CAPTCHAs require you to find 20 fire hydrants on 3 different rounds of tiles, just to fail you anyway. We have loads of data on websites keeping user's attention [1], this also seems to apply to CAPTCHAs.
Besides, I think it is now well known that AI is fully capable of solving CAPTCHAs.
[1] https://www.nngroup.com/articles/response-times-3-important-...
That's the biggest downside of modern AI, and I fear the web will only get worse because of it. If we can't figure out how to patch CAPTCHAs against bots, remote attestation will become the norm.
1. Embrace the bots and get each request (with response) super lightweight. Anything you can pre-compute, pre-compress, pre-cache is great. I've used this successfully for a small service that can scale significantly.
2. Make the cost of interacting with your service computationally expensive. For example, you could send off a problem to be solved which becomes itself a token to make one interaction. There are several problems that are computationally expensive to compute, but easy to verify.
3. Make the cost of interacting with your service require sending a significant payload - the idea being that if they launch many requests from a single network, they saturate their network. If to watch a 100MB Youtube video you had to send 1MB of random data via UDP (used to fingerprint), I suspect people abusing your service would soon find they experience dropped packets. If they struggle to send 1MB of random data, there's a good chance they would have trouble downloading 100MB of data.
4. A lot of these AIs falsify information to appear plausible. You could abuse this to ask questions, some real some false, and brief the user to answer randomly on the nonsensical questions. For example, "How many connections are there in a tripoduplex?" Something like chatGPT may see tokens for "tri" and "du" and output 3 or 2. There would also be a way to do this with images, i.e. "Select all of the cats in the image and press done", where they are all some weird trip of images.
These are just some ideas and there are obvious flaws in some of them.
I guess "Turnstile" might be what I'm running into?
Fortunately I haven’t seen the longer blocking you mention.
However relying on the near universal behaviour tracking and fingerprinting of large corporations is extreemly worrying. The better Turnstile works the more like Google's proposed Web Environment Integrity it becomes.
https://www.eff.org/deeplinks/2023/08/your-computer-should-s...
Not really. I'd say >80% of sites I visit are accessible, in that I can read the content of the page I want to look at, with javascript disabled.
Of the remainder, I can temporarily switch javascript back on for a tab with two clicks, if I want to read the contents enough, and I think the odds of that site having especially malicious javascript on it is slim. (e.g. mastodon instances)
I can also create permanent exceptions for websites in a couple of clicks too, and there are some frequently-visited sites that I have done that for.
Don't think of those who disable javascript as browsing with no javascript. Think of it as browsing with javascript disabled by default.
There are several services like Turnstile. I'm an advisor to Ocule [1] which is a similar thing, except it's a standalone service you can use regardless of your serving setup rather than being integrated into a monolithic CDN. It's a smaller company too so you can get the red carpet treatment from them, and they aren't so aggressive about blocking VPNs and privacy modes because their anti-bot JS challenges are strong enough to not need it. They're ex-reverse engineers so know a lot about what works and what doesn't. Their tech may be worth looking at if you're concerned about over-blocking.
The mention of Turnstile using proof of work/space is a bit puzzling/disappointing. That stuff doesn't work so well. There are much better ways to create invisible JS challenges. The core idea is to verify you're in the intended execution environment, and then obfuscate and randomize so effectively that the adversaries give up trying to emulate your code and just run it, which can (a) lead to detection and (b) is very slow and resource intensive for them even if they aren't detected. Proof of work/space doesn't prove much about the real execution environment.
BTW, the author asks what proof of space is. It's where you allocate a huge amount of RAM and then ensure the allocation was actually real by filling it with stuff and doing computations on it. The goal is to try and restrict per-machine parallelism by causing OOM if many threads are running in parallel, something end users won't do. Obviously it's also a brute force technique that can break cheaper devices.
It also assumes a very smooth gradient of suspiciousness. When using JS challenges though, you can be often working with binary signals (at least, that's what I was able to get). So then there's not much "maybe so/maybe no" about it. Either you detect a bot with 100% confidence, in which case you just drop the banhammer. Or you don't spot it and have to let it through.
Even down to the API endpoint and JS API names.
https://www.google.com/recaptcha/api/siteverify
https://challenges.cloudflare.com/turnstile/v0/siteverify
grecaptcha.render({ callback: function (token) { ... } });
turnstile.render({ callback: function (token) { ... } });
As soon as I saw the examples I recognised the names, I guess it's designed to be a drop in replacement?https://developers.google.com/recaptcha/docs/verify https://developers.google.com/recaptcha/docs/invisible
I feel like 400 req/s should be absorbed, especially since my 5€/month VPS can handle about 2x that, sustained. I might be missing something here, but that just doesnt seem like a peak big enough to warrant degrading user experience.
Sadly I'm often categorized by websites as "probably a bot haha get fucked", and its lost sites hundreds or thousands of $$$ worth of revenue over the years, just from me alone.
I did a non-scientific test last year, in which I ran a basic Apache instance that returned a 204 response (No Content) on $5 instance and HTTPS alone made it drop around 500 requests/second (once again, no other processing happening other than returning headers for a 204 response)[1]. My understanding at the time is that in general on lower priced virtualization options you don't get VMs which provide access to CPU instructions that speed up encryption/decryption.
[1] my comment at the time https://news.ycombinator.com/item?id=30155559
Only then will it cause significant delay. If your service drops connections at 500+ requests per second, something is off - You can configure quite the large accept backlog, like a few thousand, and you should only get an additional few ms delay per additional request.
If you're bottlenecked anywhere, it shouldnt drop connections until its like very very bad.
Thats assuming you need to access a DB or do other checks -- static content servers should start struggling on a 2-4 core system at around 10k requests per second, but that doesn't really matter for this.
That $5 VPS will also be null routed within days with the kinds of DDOS traffic sites like HIBP get.
Web developers use JavaScripts to make HTTP requests to API endpoints. The data is being consumed by a script, programmatically. Unless Javascripts are neither scripts nor programs. Good luck with that argument.
There is a W3C TAG Ethical Web Principle that states web users can conusme and display data from the web any way they like. They do not need to use, for example, a particular software or a particular web page design.
2.12 People should be able to render web content as they want
People must be able to change web pages according to their needs. For example, people should be able to install style sheets, assistive browser extensions, and blockers of unwanted content or scripts or auto-played videos. We will build features and write specifications that respect peoples' agency, and will create user agents to represent those preferences on the web user's behalf.
With respect to the JavaScripts authored by web developers and inlined/sourced in web pages, web users have no control over them short of blocking them outright. As such, arguably they are not ideal for making HTTP requests to API endpoints. Unfortunately these scripts can be, and are, used to deny web users' agency.
He wants to provide a free service, which he is not obligated to do and costs money for him to do, for individuals to check their email. He never intended for bots to check a billion emails, and obviously doesn't want to pay for that.
That should be respected, and people failing to respect it is why we see the destruction of the open web with things like remote attestation as the only way forward.
Complaints like "but the open web" or "but my exotic browser" are honestly worth nothing against a potential solution for a real, pressing issue like spam requests and bots - if we want to keep the nice things we better start coming up with alternative solutions to the real problem, because corporations will decide the future for us if we don't. Ignoring it will not work out for us.
Web services are intended for humans to use them, and all abuse also ultimately comes from humans directing computers to be abusive. Thus, rather than attaching a computer's temporary identity (an IP address) to the request, we should be attaching a human identity to it. Note here, that I don't care whether the request comes from John Smith in New York. I care about being able to ban you from the service if you are abusive, no matter how many computers you have at your disposal now or in the future.
There's lots of downsides I haven't got solutions for, such as, how do we stop sites cross referencing to reliably discover all users real identities (Google analytics would love this!), and so much more. But we'll have to give up something - if we do nothing we give up everything, maybe giving up only absolute anonymity could be preferred.
You can have a group of people sharing a block and collectively responding to abuse reports to slightly improve privacy. That at least shields you somewhat from the big tech firms.
Essentially a large number of virtual private micro-ISPs.
disclaimer: acquaintance works in the spam business. someone i tried to steer clear of, but was fascinated of what they told me.
IP reputation is a huge thing in the spam world. they pay top dollar for residential US/UK/etc IPs which they can then use for spamming others. and we're not talking one or two IPs. we're talking 100s of thousands of IPs being transacted daily globally. all for spam.
for anyone interested in seeing how low we've come, they have a huge convention in Las Vegas. one should visit it to better understand a field that has been growing like crazy but so few know it. apparently everyone is preparing for a big boom next year, with 10s of $B ready to be deployed.
The point is to not get hung up on the client as a security boundary (it isn't, can't be, won't be), but to focus on the actual harm -- excessive use of the limited resources provided.
And you have to frame your security posture as rooted in the server side throttling, heuristics, et cetera. Flipping out because the client isn't what you expected isn't going to help (determined attackers can look like "vanilla" clients), and is just going to harm the long tail of actual users who are not bots.
I mean, it's built into Safari and Cloudflare supposedly uses it for its CAPTCHAs: https://developer.apple.com/news/?id=huqjyh7k
It's not as invasive as Google's attempt to circumvent ad blockers, but it's still a remote attestation system implemented in the wild already.
Is it really built into Safari, or are you referring to the iCloud option that "privately" helps reduce Cloudflare demands?
The two companies are working together to make this an official web standard, but I haven't heard about it for a while. Maybe they're just laying low after seeing the blowback on Google's (worse) attempts at attesting devices.
Indeed. To state the obvious: the real problem are bad actors - no matter if they are nation states, cybercriminals or people running compromised devices - and their accomplices such as ISPs not responding to abuse reports.
As long as we don't get that under control (say, by threatening to cut offender countries and ISPs from the Internet and SS7 phone networks) we'll have to continue whack-a-mole'ing.
You're going to have zero horid scraping once that's in place.
Not a good choice, why not make this a hashed db that could be distributed freely? Why are bots meant to be excluded? This is bad UX, I'd want to check my emails periodically in the background, but this "anti-bot" measure is meant to make me unable to do so and then demand payment. So this is openly a fight.
> are honestly worth nothing against a potential solution for a real, pressing issue like spam requests and bots
That is not an issue for me at all. Ok, let's face it - maybe I'm just on the other side. For me scrapers are very useful because they reduce the price needed to access the data as they introduce competition. For example price comparison sites are very useful.
The other thing about "waiting" is that bots may not want to do that, or maybe some sort of deadline is sooner than such waiting would allow.
To me, requiring a unique key to be input with the search, that is created after X amount of time (both provided by the initial response and increasing exponentially for subsequent requests) seems like it could be sufficient. If the next request is sooner than X then block them for Y amount of time for attempting to bypass allowable behavior. Allow like 5-10 emails then implement the wait-based functionality so that most non-bots would be fine. After all we're talking about blocking an endpoint only ever meant to be used via the website by actual users, typically they are not trying to check thousands of emails super fast.
Programmatic? Yes. Against their wishes? Also, yes.
The guilt or bad feelings subsided after realizing I'll use more of their bandwidth how I was searching before via the browser. Using the API, I see when the last inventory update occurred, and only after this will I search their discount inventory.
(In any case, if they are on EC2, I cost them 3 to 8 cents per month, estimating high, which they have certainly recouped.)
They could (incorrectly) call you a "bot" but that's simply not true. You are human.
The amount of server resources you deplete is actually less than a modern graphical browser from an ad-sponsored team of software developers.
However perhaps you are not looking at some ads or submitting yourself to tracking or telemetry. Under the W3C web ethics principles you can consume the data from the web in any way you like.
It is your right to decide what HTTP requests you make. XHR/fetch in someone else's Javascript could be well-intentioned but too often it tries to take away some of that agency and transfer it to a web developer.
For tracking, they'll have to rely on old-fashioned "order history" and "inventory period". By these metrics, I would be baffled if they decided to block my access or close my account, even if I'm bypassing Google Tags/their AB test suite/social media analytics/ads for their other services.
Ignoring their nonconsent feels impolite, but they can dry their tears with my spent money.
I've seen many games die due to bots. If developers just straight-up allow them all the normal players quit.
> 2.12 People should be able to render web content as they want
Yes, the choice between Chrome/FF/Safari/Lynx or some other remains
Any discussion not discussing how to deal with abusers is moot
Add HTTP headers documenting the API you want them to use as a bot author will look and wonder why it's performing so badly.
The difference between human speed and computer speed is so noticeable that it can be leveraged... No high bar or complex adversarial solution needed, and the side benefit is that if a bot does persist it spreads the load.
I think a lot of these schemes rely on solving a CPU or memory based computation with JavaScript
Otherwise, great write-up! And excellent service!
And if you don't stop them, you at least slowed them down big time. This could also be enough to make the attack useless.
For example, desktop computer but clicking things without moving your mouse? Suspicious. Say you're a phone, but have desktop computer fonts installed? Suspicious. And suchlike, the precise methods are the results of a cat-and-mouse game.
If these heuristics identify your browser as suspicious, they either show you an interactive captcha, or they just refuse your request.
The war on accessibility tools remains deeply irritating.
So people employ these measure and have no clue whom they filter. Reminds me of the online shops who block you because you click on the products too fast. Congratulations, you lost a customer to keep the CPU utilization at 20%.
What's weird is that the current implementation is broken. Once you do get a CAPTCHA to fill out, the redirect will fail and you end up starting over, only to get a new CAPTCHA page. That's rather unfortunate.
Thanks you Troy for writing and sharing your experience.
This can be implemented in a way that remains transparent (albeit via JS), poses little impact on ‘good’ users, but protects against a lot of traffic patterns that may be undesirable. The cost can be scaled to match infra capability and the challenge can be a combo of the request data and time. Valid windows for that time can then be synced with cache validity which removes the need to keep tabs on any state.
For those deeper in this space. What am I missing here that prevents this from being the norm?
Meanwhile, plenty of the legitimate users are using 5 year old budget android devices, so you'd better not make that challenge too hard.
Making the browser deal with PoW challenges is only a small price to pay for what is practically a free VPN. It works great, until your entire home IP starts getting CAPTCHAs all the time, and because users don't know any better, they start blaming that darn Google/Cloudflare/Microsoft for claiming they're a bot.
This could maybe work as a legit region blocking workaround service if the VPN only allowed connecting to popular streaming sites somehow. I don't think I'd trust it though.
I think there's some merit to the idea, especially for free VPNs, but you'd need a whitelist for things like streaming services and something to prevent abuse. Of course these companies were just interested in building out botnets, but it could be done somewhat okay if the right groups of pirates and streaming customers banded together.
I'm curious to know if there have been similar open source implementations of turnstile where website operators have found ways of limiting an API call to just a browser, without captcha. Does anyone know of any?
If you want to annoy SSH brute forcing bots, endlessh is a dedicated tool for SSH connections. There are other tools for other dedicated protocols as well.
What I liked about the application-level interference is that you can do something more subtle than a block, while still feeding them nonsense, slowly.
ISPs still failing to implement modern networks and sticking with broken workarounds like CGNAT are as much to blame as the bots tainting their CGNAT IP addresses.
That said, I have a real ISP, comcast, but comcast does MITM attacks on it's users so I have to tunnel everything through various VPS I rent. Which of course means I get hit with the same cloudflare blocks. And the invisible javascript ones just go in loops no matter how many times I complete the visible side. I just close cloudflare hidden sites' tabs' now. The problem is more and more of the web is hidden behind their computational paywall... even academic journals now.
Okay, so that... that is victim blaming. If you don't want to do that, stop doing it. Besides, do you really think someone in this situation can just get a real ISP? I mean, maybe they're in a competitive market and just managed to pick a bad option, but it's unlikely.
Yes, I know at least 2 such people who chose to have a wireless ISP despite having a wired option available and affordable. One of them is even slightly technical. They usually don't run into the limits, but when they do it might be the right time to give them a nudge to get a real ISP. I was hoping that'd be the case here.
If one does not know how to limit based on number of HTTP requests, then is one really qualified to set up a "public API".
It is well-known that IP-based limits, i.e., blocklists, do not work. (While allowlists are common, e.g., academic journals, Verisign zone files, etc.)
We cannot blame the public because someone does not know how to configure a proxy to limit number of HTTP requests per IP, e.g., in a 24-hour period.
Here, the public is penalised by asking for credit card numbers because the site operator does not know how to count HTTP requests.
One would think people who do not want their details in a data breach that random people can download probably would not want to give their details to some random person whose website becomes popular. But it seems they do.
HIBP never made any sense. "Send me your private info and I will check and make sure it has not been leaked. Too late. You just leaked it to me." This sort of obvious blunder had to be fixed. Still too late for anyone who used it before HIBP was "fixed". Why place any confidence in someone who cannot spot these issues. Here, he struggles to implement an API.
Stupid websites can become popular. It happens. Many websites have sought to exploit data breaches. What better way to collect working email addresses that people care about than to let people submit them to you to check against a dump of a data breach. Unless they download the dump themselves, they have no way to confirm you actually checked anything. And if they did download the dump(s), then there is no reason to submit anything to you via an "HIBP" website.
If people are getting charged for API access, and submitted personal info to HIBP, then they should get some enforceable terms in return. If you collect peoples' information then you are liable for the damage that may result if HIBP is breached. Doubtful the API customers get any such protections.
> The widget takes responsibility for running the non-interactive challenge and returning a token
I get that a widget can return a token and then you trust the token and do the rest. But what determines whether a caller receives a tokens or not? That is, what's the actual challenge? That is, how is the distinction between bot vs. real human user actually getting decided?
As for the challenge, that seems to be within Cloudflare's implementation which then returns the token to submit with the form. HIBP then verifies the token to make sure its a match before checking for the data and sending a response.
Then I found this article[1]. Embedding a challenging with valid time in Javascript, calculating the response by obfuscated Javascript code.
[1] "Bypassing YouTube video download throttling" https://news.ycombinator.com/item?id=37117338
There are much better ways to deal with bots than to let Cloudflare further bifurcate the Internet.
Such as?
I run histograms for connections per netblock on my email servers, and even removing only the most egregious attempts at abuse almost empties my logs.
On the other hand, Cloudflare has issues with tons of less popular networks, with VPNs, with less affluent countries, with non-mainstream OSes and browsers, et cetera, all of which ends up punishing many people in ways that are completely disproportionate to the amount of abuse avoided.
It reminds me of the quote from fortune(6):
As far as we know, our computer has never had an undetected error. -- Weisert
You don't know how many people Cloudflare has marginalized because you don't see their visits.
I'm interested, do you have some resources to share?
It's just referencing an example, but creating usable tools really isn't hard.