Parsing HTML in Python with BeautifulSoup
jgc.org
jgc.org
I more or less agree with your criticism of parsing via regular expressions, but there was no need to distract us all with insulting Eric (and he should know better than to respond in kind).
First, I would have had the sort of anger and upset reaction that would have come with someone describing my code as "looking like it was written by a poor programmer". I'm sure that I would have wanted to lash out at that person.
But that reaction would have quickly been replaced with deep shame that my code had been examined by someone and found to be very poor. Once I had examined the code I would have felt very bad because I had put something out there of that quality.
It's the usual response he has; we should probably be used to it by now :)
I think poor programmer was a fair comment in the end; the code does demonstrate poor programming (rather than just poor code). On the other hand "snotty kid" struck me as too far; questioning programming ability is fair enough - even being rude about it isn't nice but at least reasonably acceptable. Personal attacks is just pram and toys throwing.
Thanks!
At first I used regexes. I would find bugs, and fix them. The bugs affected the product in delaying confirmation of content ownership, which stinks.
I noticed that the bugs didn't stop coming, so I switched to BeautifulSoup. It was faster and better. I highly recommend it for anyone using Python.
Actually, I'm a bit surprised that this article has been voted to the top of HN. It's not particularly interesting or challenging from any perspective that I can think of.
http://esr.ibiblio.org/?p=1350#comment-241727
"On the other hand, every once in while I am reminded that “programmers I’ve known” are clustered in the top 5% of ability, usually the top 1% of ability."
http://www.crummy.com/software/BeautifulSoup/3.1-problems.ht...
BeautifulSoup is also only kinda-sorta-maintained from what I recall, so it's probably better to use something else.
which lets you use CSS expressions to find what you want.
But the thrust is good -- for easy-to-describe yet tough-to-implement problems, always steal somebody else's work if you can
Apparently it's slower, though. I've only used the python version, which works very well if you use it with the old SGMLParser.
For those of you doing .NET, the HtmlAgilityPack looks very interesting. You can search using multiple paradigms such as XPath, XSLT, and Linq
http://simplehtmldom.sourceforge.net/
Its a little light on documentation, but has a familiar syntax and handles malformed HTML. I've used it in a number of projects and its been great.
Parsing malformed HTML is a nightmare, I'm impressed the programmer even gave it a shot.
There's also a lot to be said for curiosity. I'm currently building an email client in my spare time, not because I don't think there are plenty of great ones already, but because I'm interested in programming with IMAP. I'll probably open source the final result, and I don't think there's anything wrong with doing so.
And I can't believe this is ranked first on hacker news.
Clearly there's some personal politics going on (insults having been traded) but jgc does raise some good points: the original code is very tightly-coupled code to whatever Berlios outputs, and would incur a large maintainability cost.
It's also bought BeautifulSoup to my attention, which seems quite a neat utility built exactly for this sort of thing.
For these reasons it's interesting, and is why I've upvoted it.
+ The word maintenance is incorrecly used there.
In this case, the juxtaposition of the quality of the code (which was truly poor) and Raymond's opinion of himself (recall that he claimed to be a Core Linux Developer at one point) made plain the problem that many people have with the man.
Interestingly, his response to my criticism was not to say something like "You are right, but don't be nasty about it". Instead he wrote a blog posting going on about how right he is. Oddly, I share his concerns about HTML parsing, but I think he's wrong to not use an HTML parser and do everything by hand. Having done a lot of screen scraping work dealing with all the edge cases is a pain.
And, also, Raymond claims to be very thick-skinned: http://catb.org/~esr/writings/take-my-job-please.html
It looks like he didn't really understand how BeautifulSoup (or for that matter, xpath) works. For example, he seems completely unaware of the '//node' syntax which would completely sidestep the issue of encoding structure in the code.
ESR probably just made some assumptions about the capabilities of available parsing options that would be correct but for a few exceptional tools. Let's fault him for not checking his assumptions rather than name-calling about poor programming.
I actually got him to admit (in public, on his blog) that half of all he claims is untrue.