That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a trace for damages.
That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a trace for damages.
When even OpenAI is more respectful of intellectual property and website owner control than you, that's a problme.
They have three different user agent strings, anyway. The problem is obviously that you can't tell what someone does with data after you give it to them.
You also don't know that some third party isn't crawling the web with the user agent string "OAI-SearchBot" and then using the results for training or selling the data to the likes of Anthropic or OpenAI without telling them that.
In general attempting to use the user agent string for access control is not going to work and it seems like Google is being the less disingenuous party on this one by not blowing smoke.
Saying this as a European who is pro EU.
Why should they?
How is that not conflict of interest??
> Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search. https://developers.google.com/crawling/docs/crawlers-fetcher...
Exclusion from grounding does mean that your site won't get sourced in the AI overview, but I'm not sure what the click through rates are like on those.
There are of course plenty of mystery, obfuscated/camouflaged scrapers/crawlers from god knows whom. Thankfully they are easy to spot and ban, although I've definitely thought about deliberating sending them poisoned data.
Huh?
Search and AI are hand-in-hand.
They both rely on embeddings. (Unless you still do keyword-only search, but that's not as good.)
that’s just my take/feedback, take it or leave it. i won’t be engaging further as i already feel i’m going against the site guidelines with this!