The only Python one I am aware of for which code is available is: http://sourceforge.net/projects/ruya/
Edit: You might also want to take a look at http://wiki.apache.org/hadoop/AmazonEC2
Edit2: Polybot is another Python based crawler, but no code. However, the paper has some interesting ideas:
Design and Implementation of a High-Performance Distributed Web Crawler. V. Shkapenyuk and T. Suel. IEEE International Conference on Data Engineering, February 2002. http://cis.poly.edu/westlab/polybot/
http://72.14.205.104/search?q=cache:LYoRD1GTP2UJ:www.oluyede...
We've been playing with sgmlop (http://effbot.org/zone/sgmlop-index.htm) for parsing and urllib2 (http://docs.python.org/lib/module-urllib2.html) for fetching.
Also, if someone on this list wants to work on a cool web spidering project (probably using Nutch), send me a message. I'm looking for someone.