Here is the robots.txt of Google
google.com
google.com
They are disappointed when they learn we obey robots.txt, so we have them manually do searches to pull out seed lists for their 80legs crawls. It's a pain, but there's not really a way around it within the rules.
After all, google is datamining the web on an ongoing basis, in return it should willingly consent to being mined in return.
For the most part "Google" is a condensed version of the web. In theory, you could "spider" a single site (google.com) and build your own search database without having to go out and crawl the web at large.
It's an odd paradox, but I think one search site crawling another search site is not a good idea. And there is probably an infinite spider loop hiding in that process somewhere.
And incidentally, some of those searches are to find forms that it can stuff links into. My sites are constantly getting hit by botnets searching for Drupal comment forms. Luckily it uses a quirky URL format that's easy for mod_security to block.
And then, this strange loop attains self awareness... (GEB reference)
Love the 'paradox' of the infinite spider loop, hadn't considered that before.
Aside: let's say you've whipped up a spiffy new ranking algorithm, and you just need an index to launch your search engine. What's faster: crawling the web, or crawling Google? I don't think such an entrepreneur would pass on a big speed up just because of a text-file.
http://yahoo.com/robots.txt Sorry, the page you requested was not found.
Yes:
http://search.yahoo.com/robots.txt
http://groups.yahoo.com/robots.txt
http://realestate.yahoo.com/robots.txt
No: http://maps.yahoo.com/robots.txt
http://omg.yahoo.com/robots.txtEdit: must have just mistyped it, works fine.
Think of it as a shortcut.