The Simile group at MIT did something similar back around 2006. Automatic identification of collections in web pages (repeated structures), detection of fields by doing tree comparisons between the repeated structures, and fetching of subsequent pages.
The software is abandoned, but their algorithms are described in a paper:
http://people.csail.mit.edu/dfhuynh/research/papers/uist2006-augmenting-web-sites.pdf