What makes a web crawler "legit?"
When I had a site that had millions of pages, I found that sites like Baidu would crawl my site as often, if not more often than Google.
I already felt the relationship with Google was parasitic, but I looked through my logs and never found a single hit that came from Baidu and many of the other search engines that would overload my site.
I was looking at a substantial part of the site running costs going to supporting web crawlers that were not doing anything (1) to help me, or (2) to help end users (if they don't want to send Chinese users to an English-speaking web site, why crawl the site?)
So like it or not I am inclined to only allow Google and Bing in the robots.txt because Google is the only site that sends a significant amount of traffic and because Bing sends some, and Google needs some competition.
There are web crawler behaviors that are annoying: harvesting email addresses, overloading your site, etc. But how do you know who is doing something wrong with the data and who is just collecting it do do nothing with it? (Probably 95% of web crawling ex. Google.)