Stackoverflow, HTML by Regex, topmost answer
stackoverflow.com
stackoverflow.com
Second: this is not at all interesting. The person asks a sensible question and then gets some ridiculous replies.
Third: it made me remember my spat with ESR about HTML parsing: http://news.ycombinator.com/item?id=923775 and now I feel sad.
Yes, just like "Second: this is not at all interesting" is your judgment in your comment. I know we all want to keep HN different from reddit, but a little tolerance is good.
Can anyone here provide a link that makes the discussion of these typed grammars available to laymen?
Depends on your definition of 'laymen,' I guess.
Also note that PCREs actually are recursively enumerable (I think).
Regular expressions, in their original version, are equivalent to Finite-stage Machines (i.e. ... regular grammars, no recursion, no stack, no memory further than keeping the current state). You can't describe the rules of HTML with a FSM.
Perl's regular expressions contain various enhancements. Newer versions of Perl's regexes also contain direct support for recursion (but frankly, you can't call those "regular expressions" anymore).
So ... if your regex library has recursion support, then you can parse HTML (since with recursion you can parse context-free / Chomsky type-2 grammars). If it doesn't support recursion, then you can't.
Btw ... the equivalent for a context-free grammar would be a Push-down Automaton ... http://en.wikipedia.org/wiki/Pushdown_automaton , which is a FSM + a stack.
If it was XML, one might get in trouble with "<[CDATA[" sections, but regarding HTML, I don't see a real issue here.
... especially not from the pragmatic point of view. Depending on the use case and the quality of the HTML source, a "dirty" regex hack might be a far better solution and using a DOM parser.
<!DOCTYPE html>
<html>
<head>
<title>I AM YOUR DOCUMENT TITLE REPLACE ME</title>
</head>
<body>
<div>
<br id="<bl>">
</div>
</body>
</html> <br id="<br id="<br>">">
I.e. it doesn't matter what is inside the quotes.Anything that requires balanced matching is NOT parseable with standard Regular Expressions, and by not parseable I mean that you will literally have an infinite amount of bugs. Shoot me an email and I can show you the math.
Even with Perl's whiz-bang recursive not-really-regexes-regexes, it's strongly not recommended to tackle balanced matching problems like HTML or XML. It might be theoretically possible (I haven't actually checked), but your brain will leak from your ears and you probably won't get it right, no matter how smart you are.
Just reading that makes me wince.
Someone slightly not getting the joke edited it out on the basis of it being troll/rambling, then someone put it all back. The nice bit... the actual point is emphasised as a result.
Schrödinger's Cat trilogy especially