I had to do it several times to collect data from HTML pages, to put the data in a small DB for further analysis.
At the end I cam up with a shell script using following UNIX/cygwin tools:
1.) curl to download the HTML side to a file;
2.) iconv to convert the HTML to UTF8 encoding if it was in a different encoding (which was the case once);
3.) tidy -asxml -numeric -utf8
to convert the HTML page to XML;
4.) xmlstarlet (http://xmlstar.sourceforge.net) with the sel command and a bunch of XPath expressions to extract data that I needed from the page and to pipe it to other unix tools. Watch out after you have retrieved data with xmlstarlet might return XML escaped characters, so I run it through "xmlstarlet unesc"
This approach worked pretty fine for me.