Cloudflare is great at blocking users legitimate or not.
There is always going to be false positives. Even more so with a tech/privacy minded userbase which uses ad blockers, disables JS, etc.
I personally have had more issues with Google Captchas than Cloudflare Turnstile when using the dreaded Firefox / Linux combo.
Of course. Google has every incentive to make life difficult for ff users, while CF doesn't.
How does Cloudflare know whether a bot is scraping for AI purposes or any other? Or are they blocking all bots but basically adding AI to the feature because AI is the cool thing right now?
I assume a correct robots.txt, plus attempt to kill off robots.txt evaders; if you could do that reliably (the second bit is the tricky bit, of course), you could selectively exclude any bots you want.
They'll just get you through common crawl or something else innocent-looking. The urge to milk everyone's content to the last penny is just too great.
what is the difference between a scraper or crawler and an "AI scraper/crawler"?
And how does cloudflare differentiate between them?
it's explained pretty well in their blog post in another thread. They have a whitelist of bots catergorized by type, and ban everything else. How easy it is to spoof (just the UA, or UA+IP whatever) is not mentioned.
The latter doesn’t request robots.txt
Didn't they add this last year?
I'm guessing the option to block AI bots using WAF rules has been available since last year, but this simple toggle on/off option is recent (though, all it does is create a WAF rule for you, which seemingly counts towards the WAF rule limit)
Glad to see the repost. I am going to check if I have this enabled. didn't know this feature existed.
I thought so. I wish they had added it 5 years ago.