Google gets preferential permissions in robots.txt
knuckleheads.club
knuckleheads.club
Everybody (including Google) pays the same price per query/byte/whatever. CrawlCo gets ownership of the crawling IP range and is prohibited from entering any other market. GoogleCo is prohibited from doing its own crawling, but can ask (perhaps require) CrawlCo to add new crawl products which are then available to everybody.
Basically force Google to do with their crawler what Amazon voluntarily did with their datacenters.
I'd rather make http://commoncrawl.org more current, accessible and more commonly used so website publishers can see a benefit in actively supporting it to lighten the load on their servers.
When it had a monopoly, AT&T was forbidden from selling software.
If a collection like commoncrawl with bulk downloads was more useful and thus used more often, even Google would have a good reason to use it.
It's not just robots.txt, it's also cloudflare and IP-based throttling. And it is very, very commonplace: http://gigablast.com/blog.html
This doesn't change that.
"Everybody except X is allowed" rules have always been "on your honor" type restrictions. "X" can always claim to be Joe Rando.