I do know that Selenium can be used for this...but am yet to see a decent example for the same. Does anyone have any good resources/examples on JS scraping that they could share?? I would be eternally grateful.
I do know that Selenium can be used for this...but am yet to see a decent example for the same. Does anyone have any good resources/examples on JS scraping that they could share?? I would be eternally grateful.
If you can't get the data from the endpoints the javascript hits then write your scraper in javascript and have it run in a headless browser, and it's the webkit engine so most sites test their site against it heavily.
Either pull the data out of the javascript objects or trigger your extraction from the html by attaching to the events in the javascript.
I worry that it's not going to replicate a real browser accurately enough, but I'm excited to try it out a bit.
We're also trying it for integration tests, as it is much quicker than Phantom or Selenium. Even there, where we control the standards-compliant site, it isn't quite good enough yet.
Would love to see more people helping make it so, though!
first install Xvfb and pyvirtualdisplay then try this snippet https://gist.github.com/4243582
selenium is great, it can even wait for ajax requests to finish (see WebDriverWait) ..
You really don't need xvfb anymore. Use xserver-xorg-video-dummy.
I've been doing this on a site that is 100% Javascript-driven for over a year, very successfully.
It's really no different than hitting a static site with Selenium. Figuring out the proper XPath to use is often the biggest challenge: Chrome Developer tools help immensely. Also, you need to watch for delays in JS rendering, so put a lot of pauses in your scripts.
It's of course slow, so if you want to distribute it across several machines, use Selenium Grid or a queue system (SQS, Resque, etc). Setup Xvfb to run on headless Linux instances.
Here's a very old Selenium 1.0 example that scrapes the full, rendered HTML of a page. After performing a scrape like this, I would then feed the HTML into a parser such as Nokogiri http://snipplr.com/view/7906/rendered-wget-with-selenium/
[0]http://jeanphix.me/Ghost.py/
If you are planning to use phantomjs, import sh and it's commandline all the way to payday :D