The HTML5 Parsing Algorithm
webkit.org
webkit.org
and IE9: http://blogs.msdn.com/b/ie/archive/2010/03/16/html5-hardware...
Unfortunately the new parsing algorithm breaks my online banking site: http://bugzil.la/565689
This is an intentional change. Prior to the HTML5 parsing algorithm, browsers backtracked and reparsed when seeing an EOF inside a script to deal with </script> inside an inline script. This means that prior to HTML5, an accidental or maliciously forced premature end of file could change the executability properties of pieces of an HTML file.
The current magic in the spec was carefully designed and researched to permit forward-only parsing in a maximally Web-compatible way. The solution in the spec was known to break a handful of pages among lots and lots of pages listed by dmoz, but the breakage was deemed negligible.
This is probably the highest-risk change in the HTML5 parsing algorithm, and this is the first report of it breaking an "important" contemporary site. Considering that keeping the forward-only tokenization behavior is highly desirable, in the absence of evidence of more important breakage, I'm treating this as an evang issue realizing that further evidence may force us to revisit this part of the spec. But let's try to get away with forward-only parsing!
What is the logic behind this? Just because we want "forward-only parsing", we're going to try to force people not to write </script> inside of a Javascript string? That seems rather silly.
Are there any benefits to forward-only parsing besides speed? Because if this change was done in the name of performance, then we really should re-evaluate who exactly is proposing these types of changes and why.
The bug report has a perfectly reasonable workaround you can use, and you can always put your JavaScript in a separate file or escape < >. It's a good idea to do this even with the current generation of browsers, because if you don't, you sometimes get strange bugs... often bugs that are not reproducible in all browsers.
Honestly, allowing unescaped </script> inside inline javascript wasn't a good idea to begin with.
Correct me if I'm wrong, but it seems to me that in order to detect that the </script> is inside an inline script, you'll have to parse the JavaScript. Requiring every HTML5 parser to include a JavaScript parser in order to parse it correctly seems a little too much for me.
Another reason: code complexity. Backtracking makes everything a lot more complex which is the exact opposite of what they're trying to achieve with the HTML5 parsing algorithm. Less code = fewer bugs, easier to understand and easier to implement.
<\/script>
Works everywhere (it's shame so few people know about it and use uglier and invalid "</sc"+"ript>")."The first occurrence of the character sequence "</" (end-tag open delimiter) is treated as terminating the end of the element's content."
Of course browsers never implemented this correctly.
Anybody else think that if your markup language requires a 10k line parser, somebody took a wrong turn at the complexity vs simplicity fork in the road?
Just imagine how large "quirks mode" must be.
The basic idea may be simple, but by the time you have covered a reasonable fraction of real-world use cases and error conditions for a sizeable population, the resulting code is anything but.
And 10k lines still counts as "moderately simple" - 10k lines is well within scope for one developer working alone.
When I consider what I can accomplish using C vs what I can accomplish using HTML, that doesn't appear to be the expected result.
It's fairly easy to get a simpler implementation by shifting complexity from the parser to the user. But from a "total society productivity" standpoint, having a handful of developers spend a few months to implement a resilient parser is a cheap price to pay to have umpteen million people save hours of debugging for every page they create.
It would have been better for everyone if HTML had used a bog-simple Lisp-like format from day one, but without a time machine we’re stuck with these kludges of history.
All browsers that implement the HTML5 parsing algorithm
should parse HTML the same way, which means your web page
should parse the same way in Firefox 4 and the WebKit
nightly, even if it contains invalid markup.
Parsing perfect HTML5 would be easy, but one of the features of HTML5 spec is that it does that no spec did before: it defines how parsing should work exactly, even in the case of invalid markup. Also, the parser hast to deal with deprecated elements no longer in the spec (such as infamous <font>). I assums most work went into this "how to parse tag soup" part.It would have been very easy to drop all the back compat bs by saying "An HTML5 document is one that begins with the 6 bytes '<html5' and if the document is invalid, reject it. Anything else, parse however you want." Browsers that support HTML5 add text/html5 to the Accept header.
You'd still need to support HTML4 somewhere. Supporting HTML5 separately just means duplicating the common parts of the parser. The simplicity boat has already sailed.
HTML4 + HTML5 < HTML4 * HTML5Browsers are going to have a complex, ugly, recovery-enabled parser in them either way, and the effort to add HTML5 to the recovery-enabled parser isn't very big, comparatively speaking.
If you allow people to leave comments on your page or if you are serving ads, you lose a bit of control in what gets put on a particular page of yours.
In case of comments you could try and sanitize them, but with ads that's hardly possible.
I spent a few hours looking into Ragel over the weekend, and, honestly, I can't envisage how a html5 parser mostly written in Ragel would look. But I'm going to give it a crack.
HTML5 does, and most of the spec (and probably the parser) is about dealing with error cases.
Plus it's not like XML parsers are small. And HTML parsing is far more complex an affair (due to not just blowing up on error)
http://dig.csail.mit.edu/breadcrumbs/node/166
Now HTML5 has to deal with two parsers.
Breaking half or more of all the websites out there wouldn’t exactly be what I would call the “spirit of the www”.
It’s possible to do it but it’s also needlessly complex while giving you very little in return. The biggest, most problematic flaw of allowing browsers to parse invalid code, namely that different browsers might handle failure differently, is in the process of being fixed and that’s good enough, I think.
Strict parsing has been tried and it was a resounding failure. It’s not going to happen.
2. What would the point be, apart from annoying every single end user?
http://diveintomark.org/archives/2004/01/14/thought_experime...
The HTML5 parsing algorithm is what the browser vendors want and need, however. Since HTML was never specified properly - not even in HTML4, where large areas were just undefined - HTML parsers have slowly evolved through trial and error. If big sites depend on a particular behaviour in a particular browser, other browsers have tried to be bug-compatible with that browser. Browsers that failed to do so have lost market share. If a browser decided to halt on encountering anything invalid would lose all of its users instantly, since the vast majority of documents are invalid.
Parsers have slowly converged towards each other, and the HTML5 algorithm is the compromise between them that breaks the last amount of content.
In other words, the goal of the HTML5 parsing spec was never to make something nice and clean. The goal was to define how to parse the unholy mess all you web developers out there have created during the past two decades. A clear spec, a good test suite and exactly identical behaviour between browsers is a benefit for everyone.
Allowing browsers on different platforms to adjust page widths, font sizes, whether or not load images and plugins in accordance with the user's preferences is fine, but that's an entirely different issue.