However, if you want to run crawling on your own infrastructure, I recommend Scrapy (http://scrapy.org/), a python crawling framework introduced on HN last year. Scrapy solves some of the more time-consuming problems involved in writing a crawler from scratch (multiple simultaneous requests, pipelined processing, raw caching, duplicate URL filtering) and comes with nifty development and administration tools. More importantly, it has an active and helpful set of core developers and good documentation. I am comfortable with both Python and Java but I chose Scrapy over 80legs because I can crawl for free on the machines I already have and I can afford to spend more time crawling from a single IP compared to 80legs which will let me crawl much faster but isn't free. Also, with Scrapy my bot can be 'naughty' - 80legs jobs obey robots.txt and limit the crawling rate per domain.
When my crawling needs outgrow my infrastructure, I am going to look at 80legs again.