Although we do put the onus on the users to pay attention to robots.txt we realize the reality is that some amount of them won't necessarily respect it. That's one of the primary reasons behind the way we designed our crawler the way we did -- requiring people to actually specify the links it will visit (as opposed to spidering around sites following all links). Our hope at least is this requires people to put a little thought into the data they want (and where they want it from) before hitting a site.