A more interesting issue is the opposite -- many large sites have robots.txt rules that Disallow all crawlers except Google. A new search engine either 100% respects robots.txt with the result that some major properties are completely unavailable in their index, ignore robots.txt in these special cases where robots.txt configuration is unreasonable, or- crawl anything that allows Google to crawl it. None of these options are great.