Crawl a website with scrapy and store extracted results with MongoDB
isbullsh.it
isbullsh.it
import httplib2, lxml, pyquery
h = httplib2.Http(".cache")
def get(url):
resp, content = h.request( url, headers={'cache-control':'max-age=3600'})
return pyquery.PyQuery( lxml.etree.HTML(content) )
This gives you a little function that fetches any URL as a jquery-like object: pq = get("http://foo.com/bar")
checkboxes = pq('form input[type=checkbox]')
nextpage = pq('a.next').attr('href')
And of course all of the requests are cached using whatever cache headers you want, so repeated requests will load instantly as you iterate.Just something else to throw in the toolbelt ...
import requests
from lxml import etree
jquery_like_page = etree.HTML(requests.get('url').text)Keep meaning to check out more of the Pycon vids
Unfortunately, the page is literally broken and unreadable on Android ICS with Chrome :(
This is a known issue and I'm working on it! I'll try to push a "responsive" version today or tomorrow.