I'm also curious what the approach was to being able to scrape sites in a way where the code is not coupled to a known, fixed markup structure using css selectors, xpath etc.
What are your plans for the project? Will you be open sourcing the code on GitHub?