Looks cool, but I couldn't find any definitive answer about whether it obeys robots.txt or not? Just that it's upto the end-user to determine which pages get crawled.
I'm not too fussed about people crawling my sites (as you say, it's gonna happen anyway), but I do worry about certain dynamically built sections of websites that are off-limits to bots for good technical reasons.