Ask YC: Any ideas about intelligent crawlers :)
I'm thinking of creating an intelligent crawler in Python. I have a project with a friend where we'd like to crawl a few specific car-related websites, grab some of the info and look for new entries. I am wondering if there is any existing technology out there where a crawler is sent to a site and either trained (visually?) or which can understand repeating information like tables that we could use to create a proof of concept. I'd appreciate any critique of my ideas which is:
1. create a visual tool - probably windows/mac based which uses the browser to navigate a site and to highlight elements that we would like to capture, such as car name, description, price. This would also have to be able to automatically/manually work out repeating elements
2. this tool would create some kind of file (xml?) which would then be used by the main crawler to understand how to navigate the site
3. The crawler, which we'd write in python would visit the site every week to look for new information
Am I going about this the right way or does anyone have any ideas
One point, we would seek permission from the sites before crawling - it would be to their benefit as we're looking to push people their way.
Appreciate any thoughts anyone might have
All the best
John