The Perils of Web Crawling
streamhacker.com
streamhacker.com
I learned so much about the interwebs ... http, urls (their construction), html markup and why things work the way they do ... learned about threading, using queues, and finally ... really grokked OOP.
Best of all I gained a newfound respect and understanding of Googlebot and web browsers in General ... dealing with people's crazy ass html code is not. easy.
If I ever teach a class on programming ... its something I'd love to have my students attempt as a semester long (background) project.
Good times.
In particular, with screen scraping, you are trying to extract structured data from a markup language (in this case HTML) that simply doesn't guarantee the structure your looking for. With web crawling you only need the structural guarantees offered by the HTML markup (not even that, with the quality of libraries such as TagSoup or Neko).
Now, that isn't to say web crawling doesn't have its own challenges (URL canonicalization anyone?).
Wish it would happen here on HN. Seems like it's a weekly occurrence where someone announces a pet project that involves crawling the whole site. I can't help but to think this is connected to the 30+ second page loads I get here often.