AI crawlers haven't learned to play nice with websites
theregister.com
theregister.com
At 2025-02-22 02:14:52, less than a day after i set up this honeypot, 66.249.68.37 (AS15169 GOOGLE) made a request on that invisible link, using user agent string "Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/133.0.6943.53 Mobile Safari/537.36 (compatible; GoogleOther)"
So now i have hard evidence that google runs bots that ignore the rules from the robots.txt, but still parses urls from there, queries these resources anyways and follows links marked as "nofollow".
It's just for us oldies who prefer the ethical web where people respected the RFCs they have collectively written.
What will be the consequence? My risky prediction: Absolutely Nothing, for neither goggle nor any of the other miscreants running the bot nets.
And just to take a poke at the title of the original El Reg article: Do you really think the problem is that they just haven't "learned" yet how to "play nice"? I'm sure the great oracle goggle, just needs someone to poperly splain it to them.
Probably need to collect logs across multiple sites though, as some crawlers would exclude your domain once they knew.
2025-03-05 13:32:09 45.89.148.57 (AS46844 SHARKTECH) "Mozilla/5.0 (Windows; U; Windows NT 6.1; pt-BR; rv:1.9.2.18) Gecko/20110614 Firefox/3.6.18 (.NET CLR 3.5.30729)"
$db=...
$sth=$db->prepare("INSERT INTO beacon (time, addr, query, agent, referer) VALUES (FROM_UNIXTIME(?), ?, ?, ?, ?)");
$sth->execute([
$_SERVER["REQUEST_TIME"],
$_SERVER["REMOTE_ADDR"],
$_SERVER["QUERY_STRING"],
isset($_SERVER["HTTP_USER_AGENT"])?$_SERVER["HTTP_USER_AGENT"]:NULL,
isset($_SERVER["HTTP_REFERER"])?$_SERVER["HTTP_REFERER"]:NULL
]);Not a single one of them respects that request anymore. Anthropic is the worst at the moment. It starts hitting the site with an anthropic useragent. You can litterally see it hit the robots file, then it changes it's useragent to a generic browser one and carries on.
I say at the moment because they're all doing this kind of crap, they seem to take it in turns at ramping up hammering servers.
Maybe they used AI to code their agent, and it's just not that good.
AI companies have seen the prisoner's dilemma of good internet citizenship and slammed hard on the "defect" button. They plan to steal your content, put it in an attribution-removing blender, then sell it to other people. With the explicit purpose of replacing human contribution on the internet and in the workplace. After all, they represent a reality-distorting amount of capital, so why should any rules apply to them?
The robots.txt is only part of the problem, simply overloading sites is another. Maybe you don't really mind that your site is getting scraped, but you do mind that it crashes due to poorly written scrapers.
It was always understood that your scrappers should limit their load to something that a site can handle. Google has always excelled at this, I've never experienced the GoogleBot being an issue, nor the BingBot. The developers who work at the AI companies however are less talented and less careering. If their bot crashes a site, meeeh, try again later.
I don't think they DDoS sites deliberately, they just don't have the skills to fix their shitty scrapper.
If we want the problem to go away, ask AWS, Azure, GCP, Alibaba Cloud and others for a way to report bad behaviour. Make their hosting providers take action.
It doesn't matter whether they have the skills or not, because they don't care.
As long as they can siphon up enough data to increase their revenue or stock price by 0.001¢, they're perfectly content to leave small independent websites shattered and unusable under the weight of their aggressive, unrelenting scraper bots.
I'm firmly of the opinion that they should all be considered hostile, predatory companies, regardless of what supposed "value" they bring.
Train your river-guzzling models with data that's paid for or given voluntarily and then we can talk about your "value".
I have a domain parked with just a blank page that's getting 172,000 requests per month. Since it's on a free-tier static hosting, that costs me nothing. However, if that trend continues, I don't know how long we will have those free static hosting options available for everyone.
I think that it's fair to say that these bots aren't badly written. Their behavior is entirely intentional.
I think this form the article is telling:
> The Register asked Schubert about this in early January. "Funnily enough, a few days after the post went viral, all crawling stopped," he responded at the time. "Not just on the Diaspora wiki, but on my entire infrastructure. I'm not entirely sure why, but here we are."
So the owners of those crawlers certainly have the ability to stop the traffic, they're just choosing not too...until publicity makes their actions risky.
Are there packages you know of in the major language ecosystems that you can easily point at a both a robots.txt and an index.html (you need both to be compliant, most packages I've found only look at robots.txt) to get an answer on what you can and can't do at what rate?
The only one I'm aware of is a C# package which still requires you to do a lot of the heavy lifting[0].
It's much more difficult than it should be for scrapers who want to be compliant to to actually be compliant.
People are bringing up robots.txt because they want to ban AI scrapers, but if those same scrapers didn't constantly pound sites into the ground, it would be less of an issue.
They take the data and deal with the consequences later because there's none so far. I hope data poisoning see more traction as a possible countermeasure.
I guess I'm in the minority that I don't care too much about crawlers, even the use of my work for AI, but I think the current hit rates just can't be justified.
How do they know it's LLM "bots"? Did they provide any proof? Is it crawlers or automated user agents? The guy behind SH hates all "AI" and is vocal about it; it would be reassuring to see some evidence, all I've seen is them saying "it's those pesky AI bots!"