After all, google is datamining the web on an ongoing basis, in return it should willingly consent to being mined in return.
After all, google is datamining the web on an ongoing basis, in return it should willingly consent to being mined in return.
For the most part "Google" is a condensed version of the web. In theory, you could "spider" a single site (google.com) and build your own search database without having to go out and crawl the web at large.
It's an odd paradox, but I think one search site crawling another search site is not a good idea. And there is probably an infinite spider loop hiding in that process somewhere.
And then, this strange loop attains self awareness... (GEB reference)
And incidentally, some of those searches are to find forms that it can stuff links into. My sites are constantly getting hit by botnets searching for Drupal comment forms. Luckily it uses a quirky URL format that's easy for mod_security to block.
Love the 'paradox' of the infinite spider loop, hadn't considered that before.
Aside: let's say you've whipped up a spiffy new ranking algorithm, and you just need an index to launch your search engine. What's faster: crawling the web, or crawling Google? I don't think such an entrepreneur would pass on a big speed up just because of a text-file.