[1] https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
They showed no difference between search bots and archive bots. robots.txt was never for SEO alone. Sites exclude print versions so people see more ads and links to other pages. Sites exclude search pages to conserve resources. They said sites exclude large files for costs. And they can't think sites want sensitive areas like administrative pages archived.
Really Internet Archive stopped respecting robots.txt because they wanted to archive what sites didn't want them to archive. Many sites disallowed Internet Archive specifically. Many sites allowed specific bots. Many sites disallowed all bots and meant all bots. And hiding old snapshots when a new domain owner changed robots.txt was a self inflicted problem. robots.txt says what to crawl or not now. They knew all of this.
You can make a request by typing the url in chrome, or by asking an AI tool to do so. Both start from user intent, both heavily rely on complicated software to work.
It's fairly logical to assume that bots don't have an intent and users do. It's not the only available interpretation though.
> WWW Robots (also called wanderers or spiders) are programs that traverse many pages in the World Wide Web by recursively retrieving linked pages.
It's plenty logical. That doesn't make it correct.
> if they are only intended to block crawlers why aren't they called crawlers.txt instead and remove all ambiguity?
Ha. Ask HTTP Referer.
A million standards have quirks in them that we're stuck with.