Block the Bots That Feed “AI” Models by Scraping Your Website
neil-clarke.com
neil-clarke.com
If something should not be scraped I think it should be placed behind strong authentication and the AUP/ToS must state in clear and concise language explicitly what the sites content may be used for. Even then I would expect risk takers to simply ignore it and accept any legal findings to be the cost of doing business. Or perhaps I have become a bit too skeptical and jaded with time and experience. I will go run a lap.
if it's voluntary and companies will make more money by not doing it, then at least some are not going to do it.
So, from my point of view, this effort is mostly meaningless. It's not enough to get me to open my websites again, anyway.
Would it?
The primary purpose of robots.txt is not actually to lock out bots (that's why respecting it is not mandatory). It's to give the bots guidance as to which parts of your site are appropriate for them and which parts are not.
This may make the "clear intent" argument weak in court.
The standard has the keyword "Disallow", not "Avoid". I can't speak for anyone else of course, but that seems a pretty clear indicator of intent to me. By that I mean a site's stakeholders want to indicate that certain bots are disallowed from crawling a portion of their website.
I'm not saying a court wouldn't find intent signaled, I don't know, only that it's not clear-cut that it would.
What isn't known is if those sites will still have their content possibly included in training corpuses from CommonCrawl or ThePile, etc.
What isn't known is if those same sites will be also included in corpuses such as CommonCrawl or ThePile, leading to being included in training as is.
Bytedance's crawler should probably be blocked too, based on a recent mattKC video.
I assume google's training bard with their search engine crawling data, so you have an incentive to allow them.
If so, are you willing to share an example robots.txt file? What bots would be allowed? Google, Bing, ...? Those?
Any robots.txt I might suggest wouldn't necessarily be appropriate for anyone else. It's a policy that site stakeholders need to decide for themselves I'd say.