It was my understanding the urllib2 respects robots.txt automatically. I can't find much to back that up, but I really thought I read that somewhere reliable once.
That link is wrong. Python does ship with a robotparser module in the standard library that parses robots.txt files, but urllib2 does not use it out of the box. This can be easily confirmed using Wireshark or a quick glance at the source: http://hg.python.org/cpython/file/08b5e2c9112c/Lib/urllib2.p....