Web Scraping with Node.js and Chimera
deanmao.com
deanmao.com
For example, for POSTing and reading from redis/resque I wrote this (proof of concept, not what's in production):
https://gist.github.com/000037f472b72d9490a6
A few thoughts..
> There are similar "glues" like phantomjs-node that integrate phantomjs by
> spawning a process, and processing the stdout stream, but it is limited by
> what can be done via the command line of phantomjs. If you really want direct
> api access to the browser, the best way is via direct integration.
This seems like a lot of overhead on top of a phantomjs (or even just a generic webkit) worker. Substack's approach was to just put a proxy in front of a browser that injects a <script> tag into the page to boss the browser around:https://github.com/substack/schoolbus
Supposedly the actual browser client shouldn't matter, as long as your fleet of workers are up and running. I bet chimera's approach will end up with more access to npm modules in the long run compared to phantomjs.
Also, the link wasn't in the article: https://github.com/deanmao/node-chimera
For the python equivalent of this project, there's https://github.com/kanzure/pyphantomjs
I actually wrote a bookmarklet a while back that would look for CoffeeScript snippets on a page and translate them for people who find it troublesome, but didn't end up doing much with it because I didn't feel like there was much interest.
it's not an apt comparison. You can directly write javascript and run it in your browser or in node (or test it out interactively in the node REPL).
I like to play with modules in node interactively (when relevant) because its easier to see what's going on and much easier to iterate (esp. in conjunction with the .load REPL command)
$ echo 'process.stdout.write "Now with CS level #{1 + 2}!"' > foo.coffee
$ coffee foo.coffee
Now with CS level 3!
$index.js:
require('coffee-script')
module.exports = require('app.coffee')https://github.com/deanmao/node-chimera/blob/master/example....
If you want to parse the DOM for the internet at large, you need a real browser. There are simply too many sites with really bad HTML to be parsed reliably with anything else.
It's merely an integration library for the netsurf browser. It's the html parser for a real browser. I considered using the parser from other browsers like firefox or webkit, but netsurf had the fewest external dependencies.
phantomjs-node uses ridiculous hacks. Not being rude but I think the author was drunk when writing that. I ended up wasting quite some time when trying to customize it.
You do NOT need to use any hacks whatsoever to run dnode on phantom, just browserify(https://github.com/substack/node-browserify) the shoe and dnode modules and use the example code at https://github.com/substack/dnode .
Is more documentation to come?
I sorta gave up on the project seeing as nobody other than myself used it. (I have zero watchers on github)
Great work!
That is pretty cool.