Parsley: a simple language for extracting structured data from web pages
github.com
github.com
So my current thinking on the idea is reflected in https://github.com/fizx/pquery. PQuery addresses some weaknesses of Parsley by embedding the ideas in Javascript.
(1) Parsley isn't turing-complete, and many web pages are ugly, so you often have to resort to pre/post-processing in some scripting language. I never was able to get sufficient power out of a purely declarative language.
(2) Javascript environments are readily available (even in embedded form), and are more accessible than C.
(3) If your crawler already executes Javascript to render dynamic pages, then running more Javascript in that environment is pretty easy.
I guess I'm a little late to the thread (yay weekends) but I'll answer any questions people may have.
Ruby bindings: https://github.com/fizx/parsley-ruby
Python bindings: https://github.com/fizx/pyparsley
Parsley is just so good a name. It has the word parse as a kangaroo, and it evokes the image of fresh, green, edible.
I just ran
grep /usr/share/dict/words -e "^pars[^']*[^s]$"
to find the following list of words beginning with pars parse
parsec
parsed
parser
parsimony
parsing
parsley
parsnip
parson
parsonage
I think I might call my next parsing-oriented tool either parson or parsnip. :)"Parsley" is indeed a great name. I'll contribute another Parsley - a Flex framework of yore also bore that name.
After a library is named "Parsley", the list of meanings for Parsley now includes "A parsing library", and so I see it as a kangaroo for Parse.
(I changed a literal asterisk to <star> to avoid formatting.) You get a few more options if you remove the anchor at the beginning:
grep /usr/share/dict/words -e "pars[^']<star>[^s]$"
For example, 'sparse' seems like a good name for a lightweight parser library.I'd also like to give some sort of -1 for the recycled library name, though it's not a technical nit pick, just a personal one. The name of the library is mostly dominating the discussion here at the moment, and that's a shame.
Note 3/4 of the links on the main page are to not yet created wiki pages. Looking forward to it, or just writing it for myself in go :)