OpenAI, Anthropic ignoring rule that prevents bots scraping online content
businessinsider.com
businessinsider.com
There’s nothing in the linked article to distinguish between those types of requests or usage, and no evidence to suggest the firms are doing other than what they say they are.
Especially with the rise of SEO/GenAI spam content, I kind of wish media had more rigor on citing claims - rather than leaving them as bare assertions - ideally so that someone could follow the chain all the way back and see the actual evidence giving rise to (and hopefully substantiating) the claim.
Is this from first-hand investigation where you could show access logs? Is this further communication put out by "TollBit", the licensing broker, in addition to their mentioned accusations that didn't include names? Is this coming from anonymous insiders at OpenAI/Anthropic who emailed BI? Did this information materialise in front of the article's author?
> WWW Robots (also called wanderers or spiders) are programs that traverse many pages in the World Wide Web by recursively retrieving linked pages.
— http://www.robotstxt.org/orig.html
When these companies crawl the web looking for training data, they should obey.robots.txt. When a user asks them about a particular page (e.g. “summarise this page”), they should ignore robots.txt
All of these articles purport to show that these companies are not doing the former by testing the behaviour of the latter. That’s a misunderstanding of how robots.txt is supposed to work.
Which countries? (I'm genuinely curious)
> Is respecting a "no trespassing" sign also just courtesy?
In countries with right to roam, the answer is often yes. In states like California with right to access waterways, many "no trespassing" signs are unenforceable too if they block access to rivers or beaches.
The EU's AI act points to the DSM directive's text and data mining exemption, allowing for commercial data mining so long as machine-readable opt-outs are respected - robots.txt is typically taken as the established standard for this. It has generally been respected as far as I'm aware (though, this article asserts otherwise).
In the US, respecting robots.txt/<NoAI>/... could potentially be a fallback if the Fair Use defense falls through, by implied license like for website caching in Field v. Google Inc ("Google reasonably interpreted absence of meta-tags as permission to present 'Cached' links to the pages of Field's site").
However, what you do with that data is still subject to copyright law.
"learnright" would give the exclusive right to learn from a given piece of content to the author. So everybody else who wants to have their robots learn from it would need to make a deal with the author.
My expectation is that this will not happen and robots will be allowed to learn from everything that is public. Just like humans.
None of that “only Google” nonsense.
Unless you mean in a more political manner, that copyleft should be prescribed to website owners
In general, no. Copyright is the exclusive right to copy.
If you buy a copy of a bestselling book and make copies of it, the copyright holder has the right to sue you, because the government has awarded them an exclusive (limited, temporary) monopoly over creating new copies of that work.
If you buy a copy of a bestselling book and learn things from it, deface it, make origami out of it, burn it to keep warm over winter, or whatever, the copyright holder doesn’t get any say in the matter, because it’s copyright not useright.
send robots ToS clickthrough
on fail redirect.