So they sometimes hit bollards and turnstiles made for other types of code which executes HTTP requests. So they're bots basically, but better (or suitably) behaving ones.
What is the difference if I use a browser or a LLM tool (or curl, or wget, etc) to make those requests?
LLM finds out about it from me, when I ask it to go to the link.
You don’t accuse browsers of “somehow find[ing] the existence of those pages”. How does a browser know what page to visit?
The user tells it to.
If I prompt an LLM “go to example.net and summarize the page” how is that any different from me typing example.net in a browser URL bar?
I have been talking about the latter, agree the former is abusive.
Why would that be an issue?
I thought they were just machine code running on part GPU and part CPU.
There's some gray area though, and the search engine indexing in advance (not sure if they've partnered with Bing/Google/...) should still follow robots.txt.
But if I say, "Search the web for a low-carb chicken casserole recipe that takes squash and cottage cheese," then it's either going to A) send queries to a search engine like Google, in which case robots.txt already should have been respected, or B) check its own repository of information it's spidered before I asked the question, in which case it should have respected robots.txt itself.
The entire web was built on the understanding that humans generally operate browsers, and robots.txt is specifically for scenarios in which they do not.
To pretend that the automated reading of websites by AI agents is not something different…is quite a stretch.
Should I not be able to execute curl to download a webpage because the "understanding that humans generally operate browsers"?
Isn't this a bit of an oversimplification, though? Especially when the tool you're using completely alters the relationship between the content author and the reader?
I hear this argument often: "it's just another tool and we've always used tools". But would you acknowledge that some tools change the dynamics entirely?
> Should I not be able to execute curl to download a webpage because the "understanding that humans generally operate browsers"?
Executing curl to download a webpage is nothing new, and compared to a traditional browser, has about the same impact. This is still drastically different than asking an AI agent to gather information and one of the pages it happens to "read" is the one you were previously navigating to with a browser or downloading with curl.
If you're a content creator who built a site/business based on a pre-LLM understanding of the dynamics of the ecosystem, doesn't it seem reasonable to see these types of "readers" differently?
If the scale bothers you, block it, just like how you would block any other crawlers.
Other than that, we all wanted "ease-of-access" (not me though), and now we have it. It does not change anything.