titles = [x for x in soup.findAll('td','title') if x.findChildren()][:-1]
packs a lot of punch. tds = [td for td in soup.findAll('td') if td.parent.get('class') == 'blah']
In jQuery, this is more compactly written: tds = $('.blah > td')
And if you just want to look for <td>s somewhere within a .blah element, you can use tds = $('.blah td')
This is a lot less clear in BeautifulSoup: tds = [td for td in soup.findAll('td') if td.findParents(attrs={'class': 'blah'})]
(If there are better ways to write this BeautifulSoup code, please let me know)Selectors have some other benefits too - you can just go to the CSS file and grab the selector that matches what you want, and you can be reasonably sure it'll work in most cases.
I spent more than a year writing hundreds of scrapers that ran for weeks at a time. BeautifulSoup did not work out as well as lxml in practice. On extremely javascript heavy pages we used pyv8 actually.
edit: more information at http://blog.ianbicking.org/2008/12/10/lxml-an-underappreciat... the comments are useful too.
That's like saying "jQuery is overkill for just about everything, you should use plain javascript".
'scrapy startproject' creates a couple nested directories, with maybe seven files. Are you writing a scraper that you're going to run regularly? Does it need to be super robust and maintainable? Or are you writing something that you'll run once, maybe twice?