Python web scraping
jakeaustwick.me
jakeaustwick.me
I coupled scrapy with Stem to control Tor on the odd chance your IP gets blocked. Works great.
There are hundreds of proxy vendors, and other curators of open proxies if you have a tight budget, which are designed to spread out the origin of the load and do some minor masking of your traffic.
But it would make me very sad to have an exit node blocked for being involved with this.
It could just be me, but I prefer more control, and using the requests library gave me that. I ran into some obscure problems when wanting to use / change multiple proxies with Scrapy.
I think for a long running spider, then maybe scrapy is worth looking into. For simple one-off scripts / things not run very often, I'd much prefer to write something custom and not have to learn about Pipelines, Filters, etc.
First, you can just put ``request.meta["proxy"] = "http://..."`` at any point along the DownloadMiddleware before the request is actually transmitted.
Second, you could also just package up that "more control over requests" you described and make it a DownloadMiddleware. AFAIK, the first one in the chain that returns a populated Response object wins, so you could short-circuit all of the built-in downloading mechanism.
I've built private libraries on top of requests now that allow me to do everything in such a trivial amount of time, so I prefer this approach with more control.
I think if I was going to write a long running spider, I'd probably look into scrapy again beforehand.
I had the exact same opinion until a month back when I started using scrapy + scrapyd[1] for a serious scraping task. Yes, it's a long running spider which is why I decided to use Scrapy in the first place. But so far I have been really impressed with it in general and might consider it even for one-off crawling in future. IMO, pipelines etc. are worth learning. I would also recommend using the httpcache middleware during development which makes testing and debugging easy without having to download content again and again.
On a separate note, for one-off scraping, it's also worth checking whether a python script is required at all. For example, if you just need to download all images, wget with recursive download option should work pretty well in most cases. For eg. a while back I wanted to download all game of life configs (plain text files with .cells extension) from a certain site[2]. Initially I wrote a Python script, but when I learnt about wget's -r option, I replaced it with the following one-liner:
$ wget -P ./cells -nd -r -l 2 -A cells http://www.bitstorm.org/gameoflife/lexicon/
In the end I dumped my scrapy code and replaced it with a combination of celery, mechanize and pyquery. That worked much better, used less code and was much more flexible.
I deployed a small but moderately complex crawler backed by a database with Scrapy. It had some custom pipelining as well that deduplicated information (based upon a hash) and exponentially backed-off if the page hadn't changed in a while.
>>> from lxml import etree
>>> etree.LIBXML_VERSION
If you find a page in the wild that it can't handle, I'd love to know the URL.I'm planning on adding more to the article in the near future, this was just a start. I plan on it being a resource with almost everything in, so people can bookmark it for future use.
I really haven't seen lxml choking on pages, are you sure you have the latest libxml2 installed? lxml always seems to work for me. If you know of a URL where it doesn't, I'd love to see it.
Edit: so easy, in fact, that I prefer to just START with using mechanize to fetch the pages--why bother testing whether or not your downloader needs JS, cookies, a reasonable user agent, etc--just start with them.
Mechanize is awfully slow though, if you need to crawl quickly it's not asynchronous. I wouldn't want to use it for general crawling. I guess you could patch the stdlib with gevent, and try and get something working that way.
I'm still working on proper packaging so for the moment the only way to install Struct-o-miner is to clone it from https://github.com/aGHz/structominer.
I've actually built something similar to this myself, I plan on writing an article in the future with something along these lines.
Yours look pretty polished though, good job!
Thanks! It's still a work in progress, so if you have anything you'd like to see in there, I'd love to hear about it (I also welcome code contributions if you're so inclined).
Is parsing speed really an issue when you are scraping the data you are parsing?
But sometimes you won't. Sometimes you'll be assaulted with a response in a proprietary, obfuscated, or encrypted format. In situations where reverse-engineering the Javascript is unrealistic (perhaps it is equally obfuscated), I recommend Selenium[1][2] for scraping. It hooks a remote control to Firefox, Opera, Chrome, or IE, and allows you to read the data back out.
[1]: http://docs.seleniumhq.org/ [2]: http://docs.seleniumhq.org/projects/webdriver/
I mentioned selenium in the section below that, but I'll drop a note to it in the AJAX section too!
http://www.joyofdata.de/blog/using-linux-shell-web-scraping/
Okay, it's not as bold has using Headers saying "Hire Me" but I would like to emphasize that sometimes even complex tasks can be super-easy when you use the right tools. And a combination of Linux shell tools makes this task really very straightforward (literally).
Greetings from Hamburg.
Also, many times, not all functionality/features are available through the API.
Edit: By the way, without JS enabled, the code blocks on your website are basically unviewable (at least on Firefox).
The plugin architecture alone makes Scapy a hands-down winner over whatever you might dream up on your own. And with any good plugin architecture, it ships with several optional toys, you can always add your own, and there is a pretty good community (e.g. https://github.com/darkrho/scrapy-redis)
http://crawlera.com/ will start to enter into your discussion unless you have a low-volume crawler, and http://scrapinghub.com/ are the folks behind Crawlera and (AFAIK) sponsor (or actually do) the development for Scrapy.
Pyquery is an excellent module, but I think its parser is not very "Forgiving" so it might fail for some invalid/non-standard markups.
I use webscraping (https://code.google.com/p/webscraping/) + BeautifulSoup. What I like about webscraping is that it automatically creates a local cache of the page you access so you don't end up needlessly hitting the site while you are testing the scraper.
Has anyone had any success with them?
I've done extensive scraping in both Python and Ruby, as I wrote most of the scraping / crawling code at http://serpiq.com, so I can chip in.
Overall, I prefer Python. That is pretty much solely down to the requests library though, it makes everything so simple and quick. I haven't covered it in the article yet, but you can extend the Response() class easily, so you can for example add methods like images(), links(nofollow=True), etc. Overall, I just think the requests library is much more polished than anything available in Ruby.
grequests (Python) means I can make things concurrent in a matter of minutes. However in Ruby the only capable library supporting concurrent HTTP requests that I liked was Typhoeus. It just wasn't to the same standard though, and I ran across certain issues when using proxies etc.
As far as the HTML parsing goes, I don't really have any preference. Nokogiri and lxml are both equally capable.
I think they're both perfectly capable languages though, stick with what you prefer. I've been experimenting with Go lately.
I have briefly made something with Scrapy to scrape my university's websites to notify me when we get new exam results [1], and that was okay. I might be slightly abusing scrapy but it was an okay experience.
Previously I have used 'scrubyt' for Ruby to scrape things, but their homepage seems to lead to a skin-related website now. What tools/libraries do you use for scraping stuff with Ruby these days? I remember Mechanize and Nokogiri was good/decent, but it's been more than a few years since I last used Ruby.
[1] https://github.com/flexd/studweb (description in Norwegian but it's not important)
Really? I'd love for some examples for where Ruby shines when it comes to Unicode handling when dealing with web content.
I know a lot of work was done in Ruby 1.9+ to bring decent Unicode encoding support to the language, but I still see a good number of complaints/articles about issues with it.
The encoding issues I've run into with python 2 have generally been whatever framework I'm using to ingest the content took a website at face value for encoding: either it wasn't defined at all or it was defined incorrectly.
In the wild wild world of web, unless you're doing intelligent data inspection, you're just going to run into that sort of thing.
In python, that's why projects like this exist: https://github.com/LuminosoInsight/python-ftfy
They let you correct Unicode content that was decoded with the wrong encoding.
By the way, here's another problem with taking the "encoding" parameter at face value: you're opening yourself up to DoS or data corruption bugs in the case where someone tells you to use a dangerous encoding.
There have been multiple bugs found in Python's UTF-7 decoder recently, and generally they were found by people who were scraping the Web with Python. These bugs, such as [1], could cause you to write strings that corrupt your data or crash the Python interpreter. And until the latest version -- and this is possibly still the case in all versions of Python 2 -- someone could give you a gzip bomb that decompresses to petabytes of data, and tell you it's in the "gzip" encoding [2].
I'm sure there are more bugs like this out there, and that Ruby has similar lurking bugs as well, given how recently they changed their Unicode system.
Basically, you shouldn't let someone else's Web page tell you what code to run, unless it's code you're planning to run. I recommend making a short list of encodings you trust, including ASCII, UTF-8, UTF-16, ISO-8859-x, Windows-125x, and MacRoman, and maybe a few others if you're working with CJK text, and just rejecting all others.
(The x's can be filled in with digits. Don't accept UTF-7, because it's clearly horrible. And I don't have any particular reason to be suspicious of UTF-32, but I've never seen anyone seriously use it.)