Scraping made easy with jQuery and SelectorGadget (and Node.js!)
blog.dtrejo.com
blog.dtrejo.com
All tractable problems with standard solutions, but it's difficult to accept the claim that the idea of using jQuery—which is still pretty neat IMO—now makes scraping easy.
perl -MLWP::UserAgent -e 'map { $_ =~ s/<a href="([^"]+)">([^<]+)<\/a><span class="comhead">([^<]+)<.+?<span id=[^>]*>(\d+ points)/print "$1 $2 $3 $4\n" if($i++ < 3)/ge } LWP::UserAgent->new->get("http://news.ycombinator.com/")->content;'For example, if you wanted to scrape comments on HN and get a tree-like data structure, regexes would be much more difficult to write and maintain!
Incidentally, the dev tools I wrote were in javascript and they created regex that we'd test in javascript and deploy in Perl. The regex engines are identical which is why it worked.
Scraper abstraction is for people too lazy to learn regex. Get a good book on regex, and learn how to use Perl's s/// regex with the 'e' modifier. It'll change your life.
For a solid scrapper in Perl I'd use HTML::TreeBuilder / HTML::Element. Perhaps slower than regexps, but does real parsing and understands tag-soup HTML.