This is only tangentally related to the article, but on scraping HTML in general: If you're a Python user, use lxml for it. I know most content on the web will tell you to use BeautifulSoup, and lxml is something you've only heard of in connection with reading and writing XML, but lxml actually has a lovely HTML sub-package, it's faster than BeautifulSoup, and it's compatible with Python 3. I've gotten lots of good mileage out of it (and no, I'm not a developer on it :)).