Block AI bots, scrapers and crawlers with a single click
blog.cloudflare.com
blog.cloudflare.com
* Do not confuse bots with DDoS. While bot traffic may end up overwhelming your server, your DDoS SaaS will not stop that traffic unless you have some kind of bot protection enabled, for example the product described in post.
* A lot of bots announce themselves via user agents, some don't.
* If you're running an ecom shop with a lot of product pages, expect a large portion of traffic to be bots and scrapers. In our case it was upto 50%, which was surprising.
* Some bots accept cookies and these skew your product analytics.
* We enabled automatic bot protection and a of lot our third party integrations ended up being marked as bots and their traffic was blocked. We eventually turned that off.
* (EDIT) Any sophisticated self implemented bot protection isn't worth the effort for most companies out there. But I have to admit, it's very exciting to think about all the ways to block bots.
What's our current status? We've enabled monitoring to keep a look out for DDoS attempts but we're taking the hit on bot traffic. The data on our the website isn't really private info, except maybe pricing, and we're really unsure how to think about the new AI bots scraping this information. ChatGPT already gives a summary of what our company does. We don't know if that's a good thing or not. Would be happy to hear anyone's thoughts on how to think about this topic.
It's crazy; I registered a new website last month, and every day I get around ~200 visitors, for a landing page only! This site is not mentioned or advertised anywhere. The only list where you might find it is in the newly registered domains.
Well, that's one place already. Another is in the published list of new HTTPS certificates. As such, "not mentioned" doesn't hold true.
True, but it’s one of a millions and the amount of them is still crazy
> As such, "not mentioned" doesn't hold true.
I meant by me.
No registration anywhere needed, they'll find you, because you have an IP address. I've set up enough machines without any registration and some hours after they got connected, the usual suspects showed up.
And regarding bots: even if machines don't have e.g. PHP installed, they'll see oodles of attempts to access links ending in *.php. That's the place where I liked to offer randomly encrypted linux kernels for them to digest ;-)
[1] Did actually use 1K threads to do parallel TCP connections more than 20 years ago already, so 1K is an easily reachable lower limit nowadays. You'll need an ISP which allows that to be done ... or a distributed bot farm.
Edit: spelling
I assume Microsoft intends to do the same, given they have Bing and their recent stance on the matter[2].
[1] https://developers.google.com/search/docs/crawling-indexing/...
[2] https://www.businesstoday.in/technology/news/story/microsoft...
I don't have strong opinions on this either way really, I just found that a bit funny.
Edit: Publishers Target Common Crawl In Fight Over AI Training Data https://www.wired.com/story/the-fight-against-ai-comes-to-a-...
Much of the UA data, including CCBot, is from an upstream source[0]. I was torn on whether CCBot and other archival bots should be included in the configs, since these services are not AI bot scraping services. I've added an exclusion for CCBot[1] and the archival services from the recommended configs.
[0] https://darkvisitors.com/agents/ccbot
[1] https://github.com/anthmn/ai-bot-blocker/commit/ae0c2c40fd08...
If you spew garbage, your page gets de-ranked. If your page is de-ranked, the page doesn't "exist" (effectively un-findable by most of the world).
So it's the classic rock and a hard place. For me, I still strive to create high quality, useful content because I hope it helps others. I've given up concerning myself with scrapers feeding LLMs and AI.
And it's really easy to generate random images with ImageMagick, even with random text on top to feed their OCR needs.
raw scrap from 2023+ is worthless