solr might also be worth a look. http://lucene.apache.org/solr/ It's based on Lucene also but has JSON/XML connectors to alleviate some of the pain of using Nutch.
solr is a search engine (based on lucene which itself is a search engine). it won't help the OP as he wants to crawl websites.
Maybe if he wants to search through the crawled data later, he can use it.
I would recommend solr instead of lucene as it has faceted search and updates to documents which (i think ) was missing in lucene.