Better Python Scraping - Installing lxml and Beautiful Soup
wesleyzhao.com
wesleyzhao.com
Its the same as jquery, but in Python.
Working with beatiful soup quickly becomes.long.and messy and tedious
With pyquery, you get what you want with just a couple of CSS 3 selectors simplw and nice
wow android 2.2 is terrible for inputting text
That's the key advantage as I see it. If I'm scraping something it's often in a hurry and I just want it done. Not having to internalise a new API is a significant win in that respect.
I am sticking with lxml only for my scraping and html5lib to do my richtext parsing.
lxml and Scrapy use the same library on the backend, libxml2.
http://doc.scrapy.org/faq.html#how-does-scrapy-compare-to-be... http://doc.scrapy.org/topics/selectors.html
I personally really like having the structure of a framework. It lets me churn out simple project from boilerplate very quickly, and it helps keep larger projects organized.
What's the advantage of Scrapy?
I guess it looks like BeautifulSoup finally got a 4.0 alpha release out which supposedly works, but that took several years. The codebase is aged and releases are incredibly slow.
"I no longer enjoy working on Beautiful Soup..."
"Parsing is longer a competitive advantage for Beautiful Soup, and I'd be happier if I could get out of the parser business altogether."
What it really comes down it is 2x-3x smaller code, plus much faster to write it since you can just test your CSS selector in your web browser such as with FireBug before sticking it into your code.
If you find a version of ubuntu or debian where that doesn't work, file a bug!
Also, how well do these other scrapers handle Javascript? I've had to abandon some scrapes from ASP pages because they wouldn't properly handle it.
There's some overlap, but not much. I have tended to use BeautifulSoup and mechanize together. As mentioned above, BeautifulSoup is no longer being actively maintained, and I'd recommend starting with lxml in most cases. I'm still using BeautifulSoup mainly because I have most of the package memorized.
I like html5lib, which will even spit out a Beautiful Soup parse tree if that's your thing.