Parsing HTML with Regex
stackoverflow.com
stackoverflow.com
If you know the exact HTML you are working with, using regex to extract the data is in my opinion a superior way of doing it. Less lines and generally less complexity. (Such as taking the name and id of an amazon product from a single site is different from taking all the links out of any page given.)
Well, for starters, if you're not using Javascript. All four tools you mentioned are Javascript-based.
The only real reason to use regexes is when dealing with html so broken that parts of it are inaccessible through parser.
English, as any other natural language, is (at least mostly) a context free language too, but you wouldn't go around telling people that you shouldn't ever use regexen to match certain constructions in English text, right?
Whether or not there is an absolutely snug fit between CFGs formally and natural language "in the wild", so to speak, is another topic, and rather beside the point of the analogy. Context Sensitive Grammars are overly expressive, Regular Grammars much too weak, for much the same reason why they are too weak for HTML. Were there a perfect English language parser, you would not need it in order to match regular subsets of English, just as you do not need a full HTML parser in order to match regular subsets of HTML.
Because, no, you can't parse XHTML with regex. As easily shown by the pumping lemma and all that jazz.
But, there's no freaking reason why you can't tokenize an XML start tag with a regex! In fact, you'll probably find that most uses of parsers in real life have regexes to tokenize down at the level that they can handle, before using a parser on the resulting tokens for the part that actually needs to be a CFG (among other reasons, because a compiled FSM is a lot faster than even a limited LALR parser).
Looking at this specific example, we can refer to the definitions for start tags [1] and empty element tags [2] in XML, and see that all their constituent rules form a regular language (if you don't believe me, it's not too hard to go check for yourself). So, especially since the orignal question doesn't even mention 'parsing', can we all please just shut up? (unless you actually want to figure out the horrible mess necessary to define a regex from the spec :P )
I'd like to hire the minority students, since it seems that they quite literally accomplished the impossible!
There were a few students who actually passed all of our professor's test cases (at least that's what they claim).
He decides to use regex.
Now he has two problems.
Another HN user decides to repost a relevant joke from ages ago.
Now HN is going down the drain.
This comment is from Tom Christiansen of Programming Perl / Perl Cookbook fame which includes the following caveat:
So while it certainly can be done (this posting serves as an existence proof of this incontrovertible fact), that doesn’t mean it should be.
Xpath is nice for scraping HTML, though it always turns out the stuff you need is in the middle of a bunch of other text.
Chuck Norris can parse HTML with regex.
This cracked me up :)