Scrapy: New Python web crawling & scraping framework (built on Twisted)
dev.scrapy.org
dev.scrapy.org
It accounts for 98% of the time of a current job I'm running. If anyone can provide some tips it'd be much appreciated.
Also, if you don't make use of the methods that extract/unravel object trees, they may not be properly GC'd, leading to further slowdowns. I can't remember the method names exactly (might be destroy() and extract()), but they're in the docs.
For scraping specific elements in a page, the xpath/Firebug integration is a huge win. Being able to highlight an item and grab the xpath selector in Firebug saves so much time, it's not even funny.
I essentially wrote a parallelized version of scrapy which has the ability to make hundreds of requests per second, depending on available CPUs. You could never achieve that level of performance using wget.
If you're thinking about learning web development with Python, I'd suggest Django. Other Python web frameworks are TurboGears, Pylons, Web.py or Cherry.py. Django tends to have the best documentation and probably the largest community right now, however.
I am having trouble resolving the docs to the code. Is there an IRC, mailing list or forum?
here are some notes to get started from a clean install (replace your own vars)
install Ubuntu 810 apt-get update apt-get install subversion adduser --home /home/bleu bleu su bleu svn co http://svn.scrapy.org/scrapy mv scrapy-trunk scrapy ls scrapy branches tags trunk sudo root as root apt-get install python-twisted apt-get install nano su bleu source ~/.bashrc pico ~/.bashrc #add this to end of file: export PYTHONPATH=/home/bleu/scrapy/trunk
python >>> import scrapy quit() scrapy/trunk/scrapy/bin/scrapy-admin.py startproject myproject