How to parse HTML
blogs.perl.org
blogs.perl.org
These days, however, we have the HTML5 parsing algorithm, reverse-engineered from various vendors' web browsers but actually documented and implementable (still horribly complicated, but that's legacy content for you). Not only is the HTML5 parsing algorithm designed to be compatible with legacy browsers, modern browsers are replacing their old parsing code with new HTML5-compatible implementations, so parsing should be even more consistent (I know Firefox has switched to an HTML5 parser, I think IE has made a bunch of noise about it too; I don't follow WebKit all that closely, but I'd be surprised if they haven't moved towards an HTML5 parser).
From an interoperability point of view the HTML parsing algorithm is the poster child for the success of the HTML effort; there is a testsuite of several thousand tests [1] (also submitted to the W3C [2]) that has contributions from multiple browser vendors and a number of unaffiliated individuals. Although parsing isn't sexy in the way that, say, <canvas> is, getting interoperable parsing makes it much easier to create cross-browser content (at Opera we closed a huge number of site-compatibilty bugs when we landed the new algorithm).
There are also a few open-source implementations that are not tied to browsers e.g. for python (and kind of also PHP) [3], for java [4] (fun fact: the gecko C++ implementation is generated from that java implementation) and javascript [5] https://github.com/andreasgal/dom.js It would be great to see more conforming implementations for other languages, or to see libraries like libxml2 that have existing ad-hoc HTML parsers update their implementations to match the spec.
[1] http://code.google.com/p/html5lib/source/browse/#hg%2Ftestda...
[2] http://w3c-test.org/html/tests/submission/Opera/html5lib/
[3] http://code.google.com/p/html5lib/
Yep, this was mainlined in Firefox 4 (with Gecko 2.0).
> I think IE has made a bunch of noise about it too
Support is being built, it's planned for IE10.
> I don't follow WebKit all that closely, but I'd be surprised if they haven't moved towards an HTML5 parser
The HTML5 parsing algorithm has been in Webkit since the second half of 2010.
And you have not asked, but HTML5 parsing was officially released in Opera 11.6 last month.
I hope that's not related to the annoying freezes the community's been complaining about since that release...
Better, using an implementation of the HTML5 parsing algorithm means you're parsing pages the same way browsers do: Gecko (Firefox), Webkit (Chrome and Safari) and Presto (Opera) have all landed the HTML5 parsing algorithm, and Trident (IE) is in the process of getting it (the feature is planned for IE10's Trident 6.0)
The way I see it, it's ultimately about tradeoffs. I can only imagine what things would be like today if web browsers implemented a strict parsing of HTML and refused to render invalid pages. One possibility is hindered adoption of HTML by the masses. Another is that two vendors would disagree about the HTML spec and cause pages to be browser-specific. (Turns out this happened anyway :-))
Is the HTML5 spec better in terms of interop and compatibility than the previous ones? http://www.tbray.org/ongoing/When/201x/2010/02/15/HTML5
You could argue that we would have been better off new if all browsers from day one had only rendered valid html, but you need a time machine to fix that.
Some authors might use the newhtml doctype (because they have read somewhere it is better) but only test in a browser which dont support newhtml mode, so they still don't discover that the html is invalid. So we are back to square one.
The various HTML strict modes turn off "quirks" mode as well.
https://gist.github.com/1575452
This is a sanitizing HTML "parser" done in roughly 100 lines of PHP code. It does tag and attribute whitelisting, checks for protocols to prevent XSS, deals with unclosed and unopened tags, and does some other things. The biggest issue is that it's not well-factored. However, its shortness is appealing, because I understand how it works. I would have hard time trusting a library with thousands of lines of code to do input validation.
Chrome's rendering engine, and the library used to deal with parsing HTML and building a DOM tree is Webkit's Webcore[0]. V8 and Webcore are not the same thing and V8 does not provide a DOM implementation (that's webcore's job) nor does it handle any HTML parsing (that's also) webcore's job.
V8 is a javascript VM. That's it. It does not "emulate a real web browser" (let alone completely), and nor does Node.