An Introduction to Compassionate Screen Scraping
dev.lethain.com
dev.lethain.com
So "behaving like a human" on HN might result in an IP ban because /x is denied in robots.txt. And this gets really funny when you get banned randomly because of dynamic IP addresses in cloud infrastructure.
http://www.hung-truong.com/blog/2010/12/01/conditional-gets-...
Web or data scraping is what the article talks about. Still a hard problem, easily broken by minor changes to the scraped webpage, but not subject to the vagaries of OCR and computer vision or graphical interpretation problems, which is what I was expecting from the title.
I worked at a place that allowed web-based trading in the mid 90s by wrapping a web server (on a sun spark 10) around a single emulated terminal running on a SCO box.
It was called screen scraping then, too :)
I.e., technically, you should. Practically, you can't all the time.
http://stackoverflow.com/questions/3197299/urllib2-connectio...
IE if the server's log files have 100's of requests from the same IP address in successive lines then that doesn't look like human behavior.
What would have been nice for a 'best practice' document would be to show how to set the HTTP AGENT string for the crawler so that it had an identifier, version number and some contact method.
titles = [x for x in soup.findAll('td','title') if x.findChildren()][:-1]
packs a lot of punch. tds = [td for td in soup.findAll('td') if td.parent.get('class') == 'blah']
In jQuery, this is more compactly written: tds = $('.blah > td')
And if you just want to look for <td>s somewhere within a .blah element, you can use tds = $('.blah td')
This is a lot less clear in BeautifulSoup: tds = [td for td in soup.findAll('td') if td.findParents(attrs={'class': 'blah'})]
(If there are better ways to write this BeautifulSoup code, please let me know)Selectors have some other benefits too - you can just go to the CSS file and grab the selector that matches what you want, and you can be reasonably sure it'll work in most cases.
I spent more than a year writing hundreds of scrapers that ran for weeks at a time. BeautifulSoup did not work out as well as lxml in practice. On extremely javascript heavy pages we used pyv8 actually.
edit: more information at http://blog.ianbicking.org/2008/12/10/lxml-an-underappreciat... the comments are useful too.
That's like saying "jQuery is overkill for just about everything, you should use plain javascript".
'scrapy startproject' creates a couple nested directories, with maybe seven files. Are you writing a scraper that you're going to run regularly? Does it need to be super robust and maintainable? Or are you writing something that you'll run once, maybe twice?