Apple Intelligence Foundation Language Models
arxiv.org
arxiv.org
> The AFM pre-training dataset consists of a diverse dataset consists of a diverse and high quality data mixture. This includes data we have licensed from publishers, curated publicly-available or open-sourced datasets, and publicly available information crawled by our web-crawler, Applebot. We respect the right of webpages to opt out of being crawled by Applebot, using standard robots.txt directives.
The fact that you can opt-out in robots.txt only if you knew to list Applebot months (years?) ago when they started crawling is a little unimpressive.
User-agent: anthropic-ai
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: FacebookBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: GPTBot
Disallow: / User-agent: *
Disallow: /
And possibly allow-list the ones you accept. This probably won't change the fact that you may allow a vendor at one point in time, only to realise they changed their crawling use case and has been scraping data for AI training for the past 6 months (before they go public about it).It can be argued that if you are a server operator, you always know which User-agents are making requests to your resources.
I think this goes wider than Apple. Many site administrators thought robots.txt only was for (dis)allowing crawlers that created search indexes that could give them page hits in exchange, and saw nothing wrong with allowing that.
Now, many companies crawl with another goal in mind. They don’t do that in secret, but announce that in their headers.
Now the question is how much you can blame those new crawlers for doing that. Should they have made an effort to ask site administrators “hey, you allow us to crawl your site. Are you sure you mean that?” or should site administrators have noticed the new crawlers, and have started thinking whether there meant to allow them to get their content? I think it’s a bit of both, find it hard to judge who’s more to blame, but certainly think that newer crawlers are less to blame for this, as sites should by now know that this problem exists.
There’s a ton of stuff out there from them that they intentionally appear to downplay because of that fear.
Some of it is warranted though. Any little thing gets hundreds of speculative articles written and it’s hard to control a brand.