Breaking half or more of all the websites out there wouldn’t exactly be what I would call the “spirit of the www”.
It’s possible to do it but it’s also needlessly complex while giving you very little in return. The biggest, most problematic flaw of allowing browsers to parse invalid code, namely that different browsers might handle failure differently, is in the process of being fixed and that’s good enough, I think.
Strict parsing has been tried and it was a resounding failure. It’s not going to happen.
2. What would the point be, apart from annoying every single end user?
http://diveintomark.org/archives/2004/01/14/thought_experime...
The HTML5 parsing algorithm is what the browser vendors want and need, however. Since HTML was never specified properly - not even in HTML4, where large areas were just undefined - HTML parsers have slowly evolved through trial and error. If big sites depend on a particular behaviour in a particular browser, other browsers have tried to be bug-compatible with that browser. Browsers that failed to do so have lost market share. If a browser decided to halt on encountering anything invalid would lose all of its users instantly, since the vast majority of documents are invalid.
Parsers have slowly converged towards each other, and the HTML5 algorithm is the compromise between them that breaks the last amount of content.
In other words, the goal of the HTML5 parsing spec was never to make something nice and clean. The goal was to define how to parse the unholy mess all you web developers out there have created during the past two decades. A clear spec, a good test suite and exactly identical behaviour between browsers is a benefit for everyone.
Allowing browsers on different platforms to adjust page widths, font sizes, whether or not load images and plugins in accordance with the user's preferences is fine, but that's an entirely different issue.