Upton: A Web Scraping Framework
propublica.org
propublica.org
request 'http://website.com/list_of_stories.html', (err, body) ->
$ = cheerio.load(body)
callback $('#comments li a.commenter-name').map(cheerio::text)
Plus if you need to handle javascript/ajax, just replace that with jsdom/chimera with minor changes.https://github.com/mikeal/request https://github.com/sgentle/phantomjs-node
https://github.com/joeyAghion/spidey
Has a similar approach but also leaves storage (and caching) up to the end-user.
The project was paused but I'm thinking about restarting it, and I was thinking if something like diffbot or import.io could be useful for me.. any experience doing these kind of stuff?
While I'm sure this would be possible, I think a node.js scraper would probably be a better fit for that sort of a project.
Node, by itself, doesn't have full versions of a lot of the objects that Javascript on the pages would refer to (DOM, event model, etc.); phantomjs gives you all of that.
We maintain a Node to Phantom bridge for this: https://github.com/baudehlo/node-phantom-simple
What about using something like node-gir, or whatever appjs does to combine the event loops of node/v8 and chromium/v8?
Phantom has some quirks but overall it's pretty solid.
If you insist on controlling webkit through javascript, you can use gnome-seedjs. But this is problematic/annoying because there's no commonjs implementation yet.. in phantomjs you can require() node modules in the outside context. Not so much in gnome-seedjs..
Also, X means it's not actually headless, even if you're using xserver-xorg-video-dummy or xvfb. For this reason, phantomjs got rid of the X requirement a number of versions ago.
From that perspective, Upton looks pretty cool, especially the debug mode.
`select * from data.html.cssselect where url="www.yahoo.com" and css="#news a"`
Could you elaborate on the benefits of using Upton instead of this?
http://railscasts.com/episodes/190-screen-scraping-with-noko...
As well as Mechanize when working with sites that require session cookies and all that.
I'm wondering too what the advantages of Upton are?
Upton can scrape a whole set of pages. If you have a page that lists the pages you're interested in; suppose you're interested in HN commenters on front page posts, you could specify the front page URL and a selector for links to comment pages, and Upton would automatically scrape those pages and return them to you.
Upton could even write the commenter names to a CSV for you with just a filename and a CSS selector/XPath expression.
It's not stuff you couldn't do with YQL or Python/BeautifulSoup. But it's stuff that I didn't want to have to write over and over each time I wrote a new scraper.
Here's how it currently works:
1) It has a queue of domains that I have pre-processed. For the initial purposes I've restricted it to pages that I think are ecommerce based on $ signs, add to cart/basket type links etc
2) There is a visual tool that I then use to select certain parts of the page - eg price, product, image etc. I save these out as xpaths
3) Once I have done one URL I send a crawler to that domain and extract other pages that fit the profile of an ecommerce page and try to use the same mapping as number 2 above to extract the data
I have done a small video to show it in action:
http://www.screencast.com/t/riB3iiVMiSk
I'm not sure if I'm doing this the right way. If a site/page changes structure then I may have to re-map the data. I was hoping that someone would have some pointers for me in terms of any other ways to do this. Also with Javascript-heavy sites I've had some problems
If anyone has any knowledge of screen scraping, where it can be done more automatically, I'd really appreciate a steer!
One approach I found absolutely vital was to have a rewriting, caching proxy between the crawler and the upstream site. This proxy allowed me to rewrite the page content into something much simpler for the crawler to get to grips with (RSS or Atom, say). I used Celerity (http://celerity.rubyforge.org/) with a hacked-on Mechanize API to do the rewriting, which let me handle JS-heavy pages almost as easily as static HTML ones. My original inspiration for this was _why's Mousehole (the source for which is here: https://github.com/evaryont/mousehole, I've got no idea if it runs on recent Rubies).
The proxy also gives you somewhere to raise an alert if, all of a sudden, your scraping fails because of an upstream change.
One tool I always intended to make some use of, but never got round to, was Ariel: http://ariel.rubyforge.org/. It looks like it ought to be able to totally remove the need to manually extract xpaths.
- pQuery | https://metacpan.org/module/pQuery
- Mojo::UserAgent | https://metacpan.org/module/Mojo%3a%3aUserAgent
- Scrappy | https://metacpan.org/release/Scrappy
- Web::Query | https://metacpan.org/module/Web%3a%3aQuery
- Web::Magic | https://metacpan.org/module/Web%3a%3aMagic
Above are specifically for scraping but one shouldn't forget WWW::Mechanize & LWP.
My preference over last few years is with pQuery. However Web::Query is Tokuhiro's pQuery improvement and Mojo::UserAgent looks very nifty.
> Upton depends on Nokogiri, which is basically the BeautifulSoup port for Ruby.
> If you just used vanilla Nokogiri, you'd be responsible for writing code to fetch, save (maybe), debug and sew together all the pieces of your web scraper. Upton does a lot of that work for you, so you can skip the boilerplate.
Link: http://import.io/
Seems like it's not being actively developed, but again, never had a problem.
Edit: realize this is focused on single page scraping with data extraction. Could use them nicely together in fact.
I'd say "scraping" is a little more focused on extracting data from specific pages as opposed to ALL pages as in "spidering", but the two are certainly cousins if not siblings. Anemone would probably be good at the same sorts of tasks Upton is designed for (i.e. scraping data contained on multiple pages).
I would recommend querypath. Very small footprint and takes a fraction of the cpu time.
I'm using node.js with several libraries and nothing has worked so far.
Some JS in the page makes webkit die.
I think a headless Firefox is my only hope.
What tech did you use if you don't mind answering?
My Shop Data is all PHP & MySQL, with Slim, Twig & Bootstrap. The web scraping aspect is another product of mine (forgive the clunky homepage, I'm going to turn this into an API platform) - https://grabnotify.com
GrabNotify is Node.js, Mongo, PHP, Bootstrap and PhantomJS. The undocumented API allows you to create a web crawler but define a JavaScript algorithm to extract the data off the page. Some retailers have dropdowns which update stock, images, etc, so this crawler can simulate mouse events, etc. My Shop Data will supply a custom crawler algorithm for each e-commerce web site through the API.
And finally, I've written a HTML to Markdown translator to extract page descriptions but keep some formatting while being transferable to other systems that don't support HTML.
The whole legality issue of web scraping is an interesting one. I'm planning to position GrabNotify as a web crawler, page monitor and HTML -> data tool, but only if you own or have permission to scrape the original content but need a simple way to grab and monitor the HTML into data. I'm not really interested in building a business of scraping other people's content without their permission.