So who do you guys use more? Import.io or Kimono? I have heard good things about both.
So who do you guys use more? Import.io or Kimono? I have heard good things about both.
If anybody is interested, I wrote a detailed article on scraping not so long back that was well received here: http://jakeaustwick.me/python-web-scraping-resource/
Just grabbed import.io - will see if it can loginto sites and grab the data from services I am already paying thousands per month for.
EDIT:
To add some context: I pay about $3,000 per month for some monitoring services which do not have any real reportin mechanisms. So for my daily and weekly reports, I have to manually compile them and screen shot a ton of things, compose an email and send.
I want to configure a scraper to automatically grab screens of things I want regularly and email them.
I want to have a script that will grab many diff pieces of data (visual graphs, typically) and put them all into one email.
I am working with my monitoring vendors to get them to add reporting... but until that can happen - I am tired of spending a couple hours per week screen capping graphs...
* request - https://github.com/mikeal/request
* async - https://github.com/caolan/async
* cheerio - https://github.com/cheeriojs/cheerio
- Fetchbot: https://github.com/PuerkitoBio/fetchbot
Flexible, similar API to net/http (uses a Handler interface with a simple mux provided, supports middleware, etc.)
- gocrawl: https://github.com/PuerkitoBio/gocrawl
Higher-level, more framework than library.
Coupled with goquery (https://github.com/PuerkitoBio/goquery ) to scrape the dom (well, the net/html nodes), this makes custom scrapers trivial to write.
(sorry for the self-promoting comment, but this is quite on topic)
edit: polite crawlers, not scrapers.
Kind of ironic that you are saying this about web scraping ...