Does your crawler obey robots.txt rules?
Does your crawler obey robots.txt rules?
User-Agent: established_company
Allow: /some-stuff
User-Agent: *
Disallow: /
# keeps out filthy peasants
And you're either stuck following them, and not having data that would be offered up for free if you were someone else, or being a bad person and ignoring it. You don't really see the services that follow the rules.
Also, good paper on how much being on robots.txt preferred helps, which makes you a better product, which makes you more preferred...
We don't spider retailer websites. That means we don't follow links or go hardcore on building a database of products.
We hit your website:
* if someone has asked us information about a product url
* when we place an order
* weekly for regression tests
Ping us on contact@ and we're more than happy to jump on a call and describe exactly what we're doing. Most of the time we're completely un-noticeable except for the fact that you're getting more orders.
We know for sure nobody is spidering through us.