Reading it, it looks like a total hack job by a poor programmer. For example, HTML parsing it done by a bunch of regular expressions. Which include stuff like
# Yes, Berlios generated \r<BR> sequences with no \n
text = text.replace("\r<BR>", "\r\n")
# And Berlios generated doubled </TD>s
text = text.replace("</TD></TD>", "</TD>")
All the pain of maintaining these little cases could have been removed by simply running the page being walked through an actual HTML parser that produces a DOM tree. Similarly the function dehtmlize contains a limit set of HTML attributes that it will convert (e.g. it does not convert ).Also, then you get stuff like:
# First, strip out all attributes for easier parsing
text = re.sub('<TR[^>]+>', '<TR>', text, re.I)
text = re.sub('<TD[^>]+>', '<TD>', text, re.I)
text = re.sub('<tr[^>]+>', '<TR>', text, re.I)
text = re.sub('<td[^>]+>', '<TD>', text, re.I)
Why have you got four expressions there? They are all doing case insensitive matching yet there are upper and lowercase versions of each.Hmm. The rest of the code is really pretty crappy. He could have just read the thing into lxml and have done XPath queries to extract data. It would have had the advantage of making clear what parts of the page he was extracting and made maintenance easy.