RoboBrowser: Your friendly neighborhood web scraper
github.com
github.com
[1] http://search.cpan.org/~ether/WWW-Mechanize-1.75/lib/WWW/Mec...
http://search.cpan.org/~corion/WWW-Mechanize-Firefox-0.78/li...
1. More modules for Python2 than Python3. A lot of projects are forced to go down to Python2 to allow them to use the modules they want to use. 2. It's what they're familiar with. Generally the older developers like using Python2 because they know all what they're getting, no matter what strings are attached. 3. Better syntax. Apparently some people enjoy the Python2 syntax more than the Python3 one. Not using brackets for print statements seem to be the biggest plus, even though in my opinion it looks more Pythonic.
You can run headless Selenium, and speed it up by using a static Firefox instance, but even then it'll be maybe 2-3x slower than some of the others.
The only reason this is (in my opinion), better than other solutions is because you can see the physical webpage it's loading, and for the sheer ease of use that this has. You don't even really need any coding experience to get a simple test running.
You can also use Selenium tests with things like https://www.browserstack.com/automate , where TL;DR they run your selenium test on dozens of browser + platform combinations and send you the results, like screenshots and any javascript errors. If you're familiar with CI stuff, you can see how powerful this has the potential to be. It's non-trivial but very possible to run your own cluster of selenium nodes as well; check out the official Selenium Grid: http://www.seleniumhq.org/projects/grid/
[1] https://github.com/segmentio/nightmare [2] https://github.com/segmentio/daydream
So static HTML parsing only.
- Headless vs. Not - JS support vs. Not
I put a repo together with a sample script [2] for scraping leads off of a website which I will not name, but whose name rhymes with 'help'. It uses the PhantomJS browser for headless JS support. It also includes a Vagrantfile so you can avoid installing all the dependencies on your local machine.
Selenium is a much "elaborated" solution, but still, can be detected most of the time.
Disclosure: I'm DataDome co-founder. If you want to detect bad bots and scrapers on your website, don't hesitate to try out for free and to share your feedback with us https://datadome.co
- drive a browser (firefox/chrome) via already mentioned here selenium/webdriver (potentially hiding the actual browser window into a virtual X by wrapping the whole thing with xvfb-run),
- or use one of the webkit-based toolkits: phantomjs [1] or headless horseman [2].
There is also an interesting project that combines the two, i.e. it drives a Firefox (or, more precisely, slightly outdated version of Gecko) to emulate a phantomjs-compatible API. [3]
phantomjs/slimerjs are pretty popular and even have tools that run on top of them, such as casperjs [4], that geared more to automated website testing, but can be quite good at scraping or fake-APIing too.
I have a project I'm working on that will involve scraping many different websites on a daily basis. My only scraping experience so far is using cheerio[0] to scrape a single page with a 1,000 row HTML table. Should I start with something BS-based like this or should I jump straight into Scrapy? Or are there any other alternatives I should try?
Disclaimer: I'm a cofounder there
HTTP/1.1 401 Unauthorized
Server: Microsoft-IIS/8.5
WWW-Authenticate: NTLM
WWW-Authenticate: Negotiate
...
Does RoboBrowser support these kinds of protocols? I tried to get it to work with Scrapy, but it seemed non-trivial...I've done a lot of work scraping various sites and I can tell you this: basing any product on your ability to aggregate data via scraping will not work in the long run.
Eventually you will be asked not to scrape and then you'll get sued if you don't stop.
Case law is not in your favor here. See Craigslist Vs. 3Taps.
* Lets say you are Google and you want to test if the site is working correctly every day. You could code up a Python script that opens up www.google.com, searches for "facebook" and makes sure that the first result points to www.facebook.com. This script can be configured to run everyday and if someone accidentally pushes an update to the site that causes www.facebook.com to not show up as the top result, the script automatically reverts the site back to its original state. This means users continue to get best search results even if an engineer made a mistake with the ranking algorithm.
* Lets say you are Ebay and you want to make sure that the prices for products on your site is competitive with those at Amazon. You can code up a Python script which searches for some products that customers regularly buy, like an iPhone, and extract the lowest offered price at Amazon. It can then compare them with the lowest offered price of an iPhone on Ebay. If the lowest offered price on Ebay is much larger than that at Amazon, you can offer a discount. This convinces the customer that they are getting competitive offers from Ebay and stops them from writing off Ebay when they want to shop online.