So, similarly, LLM companies can see this as a signal to crawl to whole site to add to their training sets and learn from it, if the same URL is hit for a couple of times in a relatively short time period.
I mean, currently the AI request comes from the datacenter running the AI, but eventually one of two things will happen.
AI models will get small/fast enough to run on user hardware and use the users resources: End result? You lose. The user will set their own headers and sites will play the impossible game of identifying AI.
AI sites will figure out how to route the requests via any number of potential methods so the requests appear to come from the user anyway: End result? You lose. The sites attempting to block will play the cat and mouse game of figuring out what is AI or not AI.
Note, this doesn't mean AI blocking isn't worth doing, if nothing else to reduce load on the servers. It's just not a long term winning strategy.
You may not be able to stop AIs from crawling web sites through technological means. But you can confiscate all the resources of the company that owns the AI.
Where do we stop here? at "please drink a verification can and maintain eye contact at all times"?
This is ridiculous and plain evil.
Absolutely nothing has to obey robots.txt. It’s a politeness guideline for crawlers, not a rule, and anyone expecting bots to universally respect it is misunderstanding its purpose.
And absolutely no one needs to reply to every random request from an unknown source.
robots.txt is the POLITE way of telling a crawler, or other automated system, to get lost. And as is so often the case, there is a much less polite way to do that, which is to block them.
So, the way I see it, crawlers and other automated systems have 2 options: They can honor the polite way of doing things, or they can get their packets dropped by the firewall.
It's a search engine. You 'ask it to read the web' just like you asked Google to, except Google used to actually give the website traffic.
I appreciate the concept of an AI User-agent, but without a business model that pays for the content creation, this is just going to lead to the death of anonymously accessible content.
Edit: Maybe that's fine, maybe that's bad. Maybe new models will emerge and things will reshape. But I'm just supporting the case that AI agents will pressure the current "free" content economy.
Is that a world we actually want?
Did all those old sites have “business models”? What did the web feel like back then?
(This is rhetorical - I had niche hobby sites back then, in the same way some people put out free zines, and wouldn’t give a damn about today’s AI agents so long as they were respectful.
The web was better back then, and I believe AI slop and agents brings us closer to full circle)
"What," he was asked, "is the business model for free WiFi?"
"What," he retorted, "is the business model for free washrooms?"
Many of these sites business model was simply "don't cost too much". The moment the web got big a lot of these sites died. Now add DDOS for fun and profit became a thing, most people moved to huge advertising based providers/hosters (think FB).
Simply put, we're never getting the old web back. Now, we may get something new, but it will be different and still far more commercial.
https://abcnews.go.com/Business/story?id=88041&page=1
As were punch the monkey and similar banner ads
https://www.computerworld.com/article/1360466/i-refuse-to-pu...
When was the great age of the web that wasn’t inundated with ads and SEO?
It was really easy on old school search engines like Altavista.
As for funding "content creation" itself, you have patronage.
Doesn't matter. The robots-exclusion-standard is not just about webcrawlers. A `robots.txt` can list arbitrary UserAgents.
Of course, an AI with automated websearch could ignore that, as can webcrawlers.
If they chose do that, then at some point, some server admins might, (again, same as with non-compliant webcrawlers), use more drastic measures to reduce the load, by simply blocking these accesses.
For that reason alone, it will pay off to comply with established standards in the long run.
You’d already be blocking me as I’d guess I now search via AI >90% of the time between perplexity, chatgpt, deep research, and google search AI.
If that happens a big majority of websites will go bankrupt and won't exist anymore to be searched. Problem solved!
I think that is funny considering it is likely going to have the exact opposite effect.
Low effort blog spam is cheap to make. And it is often part of content marketing strategies where brand visibility is all that matters, so not much harm if the viability is directly on your site or in an AI chatbit interface.
Quality content on the other hand is hard to make. And there are two groups of people who make such content:
1. individuals or small groups that like to share for the sake of sharing. They likely won’t care about the AI crawlers stealing their content, although I think there is a big overlap between people who still run blogs and those who dislike AI.
2. small organizations that are dedicated to one specific topic and are often largely ad financed. These organizations would likely stop to exist in such an AI search dominated world.
> Especially since website hosting is close to being free these days.
It is under specific circumstances. The problem is that those AI crawlers don’t check by once in a while like Google does but instead they hit the site very frequently. For a static site this won’t be much of an issue except for maybe bandwidth. For more complex sites like - say - the GitLab instances for OSS projects, reality paints a different picture
Another point you're missing is that there's a 3rd group of people sharing content: experts who are there to establish their expertise. Small companies and individuals generate the highest quality content these days. I work on a blog for our SAAS company and it has been a great success in terms of organic growth (even people coming from LLMs) and to simply establish authority and signal expertise in the field. I can imagine a future where this is majority of expert content on the web and it seems quite sustainable imo.
If that's websites want, they should have that option.
robots.txt is not a security mechanism, and it doesn’t “control bots.” It’s a voluntary convention mainly followed by well behaved search engine crawlers like Google and ignored by everything else.
If you’re relying on robots.txt to prevent access from non human users, you’re fundamentally misunderstanding its purpose. It’s a polite request to crawlers, not an enforcement mechanism against any and all forms of automated access.
One of my websites that gets a decent amount of traffic has pretty close to a 1-1 ratio of Googlebot accesses compared to real user traffic referred from Google. As a webmaster I'm happy with this and continue to allow Google to access the site.
If ChatGPT is giving my website a ratio of 100 bot accesses (or more) compared to 1 actual user sent to my site, I very much should have to right to decline their access.
are you trying to collect ad revenue from the actual users? otherwise a chatbot reading your page because it found it by searching google and then relaying the info, with a link, to the user who asked for it seems reasonable
- Ability to prevent their crawlers from accessing URLs via robots.txt
- Ability to prevent a page from being indexed on the internet (noindex tag)
- Ability to remove existing pages that you don't want indexed (webmaster tools)
- Ability to remove an entire domain from the search engine (webmaster tools)
It is really impolite for the AI chatbots to go around and flout all these existing conventions because they know that webmasters would restrict their access because it's much less beneficial than it is for existing search engines.
In the long run, all this is going to lead to is more anti-bot countermeasures, more content behind logins (which can have legally binding anti-AI access restrictions) and less new original content. The victim will be all humans who aren't using a chatbot to slightly benefit the ones who are.
And again, I'm not suggesting that AI chatbots should not be allowed to load webpages, just that webmasters should be able to opt out of it.
> It is really impolite for the AI chatbots to go around and flout all these existing conventions because they know that webmasters would restrict their access because it's much less beneficial than it is for existing search engines.
I agree with you about the long run effects on the internet at large, but I still don't understand the horse you have in it personally. I read you as saying (1) it's less about ad revenue than content control, but (2) content control is based on analysis of benefits, i.e. ad revenue?
Technically you don’t, but there are still laws that affect what you can legally do when accessing the web. Beyond the copyright issues that have been outlined by people a lot more qualified than me, I think you could also make the point that AI crawlers actively cause direct and indirect financial harm.
The agent should respect robots.txt no matter who is using the Robot.
robots.txt is intended to control recursive fetches. It is not intended to block any and all access.
You can test this out using wget. Fetch a URL with wget. You will see that it only fetches that URL. Now pass it the --recursive flag. It will now fetch that URL, parse the links, fetch robots.txt, then fetch the permitted links. And so on.
wget respects robots.txt. But it doesn’t even bother looking at it if it’s only fetching a single URL because it isn’t acting recursively, so robots.txt does not apply.
The same applies to Claude. Whatever search index they are using, the crawler for that search index needs to respect robots.txt because it’s acting recursively. But when the user asks the LLM to look at web results, it’s just getting a single set of URLs from that index and fetching them – assuming it’s even doing that and not using a cached version. It’s not acting recursively, so robots.txt does not apply.
I know a lot of people want to block any and all AI fetches from their sites, but robots.txt is the wrong mechanism if you want to do that. It’s simply not designed to do that. It is only designed for crawlers, i.e. software that automatically fetches links recursively.
Without recursive crawling, it will not possible for a engine to know what are valid urls[1]. They will otherwise either have to brute-force say HEAD calls for all/common string combinations and see if they return 404s or more realistically have to crawl the site to "discover" pages.
The issue of summarizing specific a URL on demand is a different problem[2] and not related to issue at hand of search tools doing crawling at scale and depriving all traffic
Robots.txt does absolutely apply to LLMs engines and search engines equally. All types of engines create indices of some nature (RAG, Inverted Index whatever) by crawling, sometimes LLM enginers have been very aggressive without respecting robots.txt limits, as many webmasters have reported over the last couple of years.
---
[1] Unless published in sitemap.xml of course.
[2] You need to have the unique URL to ask the llm to summarize in the first place, which means you likely visited the page already, while someone sharing a link with you and a tool automatically summarizing the page deprives the webmaster of impressions and thus ad revenue or sales.
This is common usage pattern in messaging apps from Slack to iMessages and been so for a decade or more, also in news aggregators to social media sites, and webmasters have managed to live with this one way or another already.
No it doesn’t. It politely requests to crawlers that they do not, and if said crawlers choose to honour it than those specific crawlers will not crawl. That’s it. It can and is ignored without penalty or enforcement.
It’s like suggesting that putting a sign in your front yard saying “please don’t rob my house” prevents burglaries.
> Robots.txt does absolutely apply to LLMs engines and search engines equally
No it doesn’t because again, it’s a request system. It applies only to whatever chooses to pay attention to it, and further, decides to abide by any request within it which there is no requirement to do.
From google themselves:
“The instructions in robots.txt files CANNOT ENFORCE crawler behavior to your site; it's up to the crawler to obey them.”
And as already pointed out, there is no requirement a crawler follow them, let alone anything else.
If you want to control access, and you’re using robots.txt, you’ve no idea what you’re doing and probably shouldn’t be in charge of doing it.
It does not. It applies to whatever crawler built the search index the LLM accesses, and it would apply to an AI agent using an LLM to work recursively, but it does not apply to the LLM itself or the feature being discussed here.
The rest of your comment seems to just be repeating what I already said:
> Whatever search index they are using, the crawler for that search index needs to respect robots.txt because it’s acting recursively. But when the user asks the LLM to look at web results, it’s just getting a single set of URLs from that index and fetching them – assuming it’s even doing that and not using a cached version. It’s not acting recursively, so robots.txt does not apply.
There is a difference between an LLM, an index that it consults, and the crawler that builds that index, and I was drawing that distinction. You can’t just lump an LLM into the same category, because it’s doing a different thing.
Yes it does. I am the one controlling robots.txt on my server. I can put whatever user agent I want into my robots.txt, and I can block as much of my page as I want to it.
People can argue semantics as much as they want...in the end, site admins decide what's in robots.txt and what isn't.
And if people believe they can just ignore them, they are right, they can. But they are gonna find it rather difficult to ignore when fail2ban starts dropping their packets with no reply ;-)
(I noticed Claude, OpenAI and a couple of others whose names were less familiar to me.)
https://github.com/bluesky-social/proposals/tree/main/0008-u...
So they sometimes hit bollards and turnstiles made for other types of code which executes HTTP requests. So they're bots basically, but better (or suitably) behaving ones.
What is the difference if I use a browser or a LLM tool (or curl, or wget, etc) to make those requests?
LLM finds out about it from me, when I ask it to go to the link.
You don’t accuse browsers of “somehow find[ing] the existence of those pages”. How does a browser know what page to visit?
The user tells it to.
If I prompt an LLM “go to example.net and summarize the page” how is that any different from me typing example.net in a browser URL bar?
I have been talking about the latter, agree the former is abusive.
Why would that be an issue?
The entire web was built on the understanding that humans generally operate browsers, and robots.txt is specifically for scenarios in which they do not.
To pretend that the automated reading of websites by AI agents is not something different…is quite a stretch.
Should I not be able to execute curl to download a webpage because the "understanding that humans generally operate browsers"?
Isn't this a bit of an oversimplification, though? Especially when the tool you're using completely alters the relationship between the content author and the reader?
I hear this argument often: "it's just another tool and we've always used tools". But would you acknowledge that some tools change the dynamics entirely?
> Should I not be able to execute curl to download a webpage because the "understanding that humans generally operate browsers"?
Executing curl to download a webpage is nothing new, and compared to a traditional browser, has about the same impact. This is still drastically different than asking an AI agent to gather information and one of the pages it happens to "read" is the one you were previously navigating to with a browser or downloading with curl.
If you're a content creator who built a site/business based on a pre-LLM understanding of the dynamics of the ecosystem, doesn't it seem reasonable to see these types of "readers" differently?
If the scale bothers you, block it, just like how you would block any other crawlers.
Other than that, we all wanted "ease-of-access" (not me though), and now we have it. It does not change anything.
I thought they were just machine code running on part GPU and part CPU.
There's some gray area though, and the search engine indexing in advance (not sure if they've partnered with Bing/Google/...) should still follow robots.txt.
But if I say, "Search the web for a low-carb chicken casserole recipe that takes squash and cottage cheese," then it's either going to A) send queries to a search engine like Google, in which case robots.txt already should have been respected, or B) check its own repository of information it's spidered before I asked the question, in which case it should have respected robots.txt itself.
[1] https://blog.google/technology/ai/an-update-on-web-publisher...
they're literally asking to break laws to train AI for national security. A sentence in a press release from 2 years ago is worthless... look at what they're actually doing
so not seem to or apparently but matter of fact like. robots.txt works for the intended audience
I'm just not sure if legal would love me doing that on our corporate servers...