Ask HN: How to aggregate product info from other websites
I have zero experience web scraping of any kind so any direction would be helpful. Before I start digging I figured HNers may have some invaluable advice.
I have zero experience web scraping of any kind so any direction would be helpful. Before I start digging I figured HNers may have some invaluable advice.
http://www.igvita.com/2007/02/04/ruby-screen-scraper-in-60-s...
Python has SGML SAX parser and since HTML is SGML it can be used. Better than regexps any day.
Python's http client library also supports cookies so that you can pretend to have a "session" with your target website.
EDIT: the libraries are urllib2, sgmllib, cookielib
What experiences has everyone else had?
If you're wondering why, well, consider this script that "learns" how to scrape Google results (from one supplied example of output data):
google_data = Scrubyt::Extractor.define do
fetch 'http://www.google.com/ncr'
fill_textfield 'q', 'ruby'
submit
link "Ruby Programming Language" do
url "href", :type => :attribute
end
next_page "Next", :limit => 2
end
puts google_data.to_xml
Reads almost like English in the scraping part!Example: http://developer.yahoo.com/shopping/