Can regular expressions parse HTML or not?
johndcook.com
johndcook.com
Case in point: A while back I ran into a nasty bug in our software which would cause it to crash with a stack overflow exception. It turned out to be that the problem was in how HTML Agility Pack would respond to certain kinds of particularly poorly-formed HTML documents.
Now, we had chosen to use HTML Agility Pack in part because of all the comments to the effect of "NOOOOOO don't use regex, use Agility Pack!" complete with links to a certain famous Stack Overflow answer that are littered around the internet. But in analzying the situation in more detail, we discovered a few things: First, we were just stripping tags, so our parsing needs were very minor compared to what HTML Agility pack does. Second, generating a full-on parse tree for an HTML document was kind of harming us, in that doing so takes both time and RAM, two things we were trying to be light on. Third, we could afford to be fairly tolerant of a missed tag or some lost content, but the parser giving up was a minor tragedy. Noisy results were far better than no results for our purposes.
The upshot of all this being, for our "HTML parsing" needs, regular expressions turned out to be exactly the right tool for the job. Following that realization, it was the work of only one afternoon to fix the bug, and as a side benefit dramatically improve both the engine's performance and the quality of its output.
Long story short, what passes for common wisdom on the Internet is no substitute for knowing what you're doing.
One easy way to get an idea of how complicated this can be is checking out the parsing algorithm in the HTML5 spec [1] (section 12).
To start the complexity, HTML is not XML and not all tags are created equal. Void elements like <img> and <br> don't have matching close tags and contents for things like <textarea> and <script> are treated as text (you can't nest more tags inside). Then you have the issue of javascript being able to call document.write at HTML-parsing times, something you need to hope doesn't happen or parsing becomes intractable. Finally, There is a vast number of rules and special cases that kick off when you have missing or mis-nested tags or tags being put in the wrong places.
[1] http://www.whatwg.org/specs/web-apps/current-work/multipage/
For example, WebKit sets a maximum DOM nesting limit of 512 [1], and most other browsers have a limit as well. A context-free language restricted to a finite maximum production depth becomes a regular language, so WebKit's parsing could be done by an NFA or DFA, and the parser could be encoded as a regular expression if desired. But the only plausible way to write such an expression correctly would be to mechanically "compile" it from a grammar. So you're going to need the grammar anyway, at which point you might as well use it directly.
[1] http://trac.webkit.org/browser/trunk/Source/WebCore/page/Set...
I just find it interesting that the efficacy of regular expressions can be framed as a computer science question, a practical question, and a statistical question.
A knee-jerk reaction to 'HTML' and 'regular expression' being used in the same sentence is like someone seeing 'goto' without understanding the context and shouting "goto considered harmful!"
I recently use some combination of perl's split() and regexes to trivially pull all links from a piece of markup. I had to suppress the "Don't parse HTML with regular expressions!" voices echoing in my head the whole time I was writing it. I'm OK with the code now, of course.
Using xpath (or just about anything else) would be smashing.
The likelihood of the markup changing to the point where the code breaks is unlikely, the target pages are open source mirrors exposed over HTTP. It even winds up being nicer code than the FTP handler.
"Don't parse HTML with regex" is good general advice, but no more an absolute than avoiding goto statements.
Or something as simple as "break 2;"
If you need to collect a number of links from the web using a spider, and you don't really mind if you get them all or not, then regex is perfectly fine. You could do a regex to match the first href="" following "<a" and stopping when it hits a "/>". This will give you at least 75% of the links on the internet, and it's done it in far less processing power or ram usage as a full html parser.
As always, a chainsaw is not necessarily better than a butter knife for cutting down trees, if the 'tree' in question is just a 5cm shoot. In fact, using a chainsaw for that would be mighty funny.
Which is why I often recommend "parsing" HTML with a regex. Particularly on throwaway projects. The overhead of a real parser is a waste.
And if the complexity grows - the sucker will need a rewrite anyhow.
But then, if the HTML has syntax problems or contains inconsistent syntax, then no approach is entirely reliable, which means it's not about regular expressions any more.
Post adds nothing of substance to the following:
"Well-formed HTML is context-free. So you can match it using regular expressions."
"But most HTML you see in the wild is not well-formed. And just because you can, doesn’t mean that you should."
which wasn't even written by him.
READ THIS FIRST: Need help? 1) Language/platform. 2) Sample string. 3) Desired result. 4) Your attempt. | Do NOT use RegEx to parse HTML! | Do NOT tell us a RegEx doesn't work if you aren't testing it in the language you asked about! | Intro: http://bit.ly/XrnV | Home: http://bit.ly/KEo1Gx | FAQ: http://bit.ly/L949Mk | Quiz: ? quiz | Regex website: http://regex101.com/
Note the part "Do NOT use RegEx to parse HTML!" ...
You can use regex to to parse HTML, but it's a verry bad idea. Basicly you would nee to reimplement the HTML definition in regex, otherwise your regex will screw up.
See http://stackoverflow.com/questions/590747/using-regular-expr... for details. Also
http://stackoverflow.com/questions/5175840/is-html-a-context...
And then you have attempts to recover when HTML is ill-formed: http://en.wikipedia.org/wiki/Tag_soup