Beating Google With CouchDB, Celery and Whoosh (Part 1)
andrewwilkinson.wordpress.com
andrewwilkinson.wordpress.com
For those considering something like this, you might want to consider using scrapy, a Python web crawler, instead of rolling your own crawler.
I remember when I was looking for something like this a year ago and found that project. I've used it for a few things and it does a nice job abstracting away most of the core scraping architecture, but leaving room to be extended as necessary. After playing with it for a few weeks, makes you realize just how easy it is to grab a bunch of data from the web if you need it.