The true power of regular expressions
nikic.github.com
nikic.github.com
While you could match well-formed HTML, the inevitable follow-up will be "how do I match parsable, but not well-formed HTML"? There is a reason that XHTML was deeply unloved.
(also... The <center> cannot hold it is too late. .. .ZA̡͊͠͝LGΌ ISͮ̂҉̯͈͕̹̘̱ TO͇̹̺ͅƝ̴ȳ̳ TH̘Ë͖́̉ ͠P̯͍̭O̚N̐Y̡ H̸̡̪̯ͨ͊̽̅̾̎Ȩ̬̩̾͛ͪ̈́̀́͘ ̶̧̨̱̹̭̯ͧ̾ͬC̷̙̲̝͖ͭ̏ͥͮ͟Oͮ͏̮̪̝͍M̲̖͊̒ͪͩͬ̚̚͜Ȇ̴̟̟͙̞ͩ͌͝S̨̥̫͎̭ͯ̿̔̀ͅ)
http://stackoverflow.com/questions/1732348/regex-match-open-...
If you ever find yourself constructing a recursively enumerable grammar (or even a CFG) using a regular expressions - whether PCREs or any other variant - you should ask yourself why you aren't using a parser generator or a proper tool for creating a compiler front-end.
I hope people don't miss the author's closing point, which is the most important part:
> But don’t forget: Just because you can, doesn’t mean that you should. Processing HTML with regular expressions is a really bad idea in some cases. In other cases it’s probably the best thing to do.
I disagree that there are cases in which it's probably the best thing to do. Most languages support XPATH/CSS selectors/etc., which are much better tools for matching arbitrary HTML patterns. I'm guilty of conjuring up a regex to scrape image links every now and then, but you should really only do that when your domain of expected input data is far more restricted than the actual CFG that you're dealing with.
Any idea how to supersede that temptation? That is, what's wrong with the HTML-processing tool/sublanguage that makes regexes attractive instead? (For me, it's that I don't need to scrape HTML quite often enough to want to learn to use what's available. I'm probably irrationally lazy.)
My recommendation is to find a library that provides jQuery/CSS selector style syntax and semantics, and then suddenly it is a lot easier to deal with the document. For example for Python there is soupselect or cssselect.
Amusingly the latter shows that the selector "div.content" translates into the XPATH "descendant-or-self::div[@class and contains(concat(' ', normalize-space(@class), ' '), ' content ')]".
< p id='some_id'>
It would be ideal to use regular expression find and replace to look for: < ([a-z])
and replace with: <$1
Of course, be sure to review every replacement to make sure it isn't part of javascript or something like that.IMO it'd be faster than writing a script and then running it against the file.
So, the article's message boils down to "if you make your language more powerful, you can do more with it. The PRCE language is so powerful that you can do the following with it: ..."
It's all pretty much over my head, so I couldn't figure out if StoneCypher was trolling or if he had a real point.
There's actually a mathematical proof out there that no regular expression engine will safely extract from all broken HTML
seems like it has to be false unless it's much more qualified.>> You cannot parse HTML with regular expressions, because HTML isn’t regular. Use an XML parser instead.
> This statement - in the context of the question - is somewhere between very misleading and outright wrong.
But nope, after disappointment and going back to the beginning, he says he's not talking about HTML:
> What I’ll try to demonstrate in this article is how powerful modern regular expressions really are.
And not even a warning about how easy it is to make really terrible regex.