Idiosyncrasies of the HTML Parser
htmlparser.info
htmlparser.info
> IE6 had an interesting HTML parser. It did not necessarily produce a tree; rather it would produce a graph, to more faithfully preserve author intent. Ill-formed markup, e.g., <em><p></em></p>, would result in an ill-formed DOM. This could cause scripts to go into infinite loops by just trying to iterate over the DOM.
Huh, didn’t know that. It’s the sort of thing that makes me wonder, “hmm, was that what was happening that time all those years ago?”
> The problem is I now have to determine which of these four options to make the other three browsers implement (that is, which do I put in the spec). What do you think is the most likely to be accepted by the others? As a reminder, the options are incestual elements that can be their own uncles, elements who have secret lives in the rendering engine, elements that change their mind about who their parents are half-way through their childhood, and quantum elements whose parents change depending on whether you observe their birth or not.
That whole section is great reading, showing just a little bit of the madness that preceded the specification of HTML parsing, and how monumental a job it was.
I think it was always a graph. I remember reading DOM implementation years ago, where a node will have siblings references as well making it graph DS.
graph doesn't mean multiple parents all the time. if you have references of other nodes in current node data(not just as children) as long as you can travel across Data structure, enough to say it as graph structure.
https://blogs.windows.com/msedgedev/2017/04/19/modernizing-d...
The decision of some website/internet property owners to decide to 'wall' themselves off has nothing to do with the desire to create well formed documents. They've done this walling off already
"Be strict in what you accept and what you return." is a better methodology.
You're right, look at what a dumpster fire HTML and general HTTP servers are under the hood as a result of building protocols on top of sloppy text formats.
our parser back then would never produce ill-formed DOM, and in this case it would have produced DOM as if for <em><p></p></em> - i.e. it would have inserted </p> upon meeting </em> to match the <p> and would have dropped the unmatched as result </p> after </em>.
The parser I know best would store a special pointer from the <em> element to where it ends, i.e. the text element between <p> and </em> and keep track of that state to end <em> rendering after passing the closing tag.
<em><p></em></p> would have become a linked list of four nodes <em> --> <p> --> </em> --> </p>
To efficiently find siblings there are forward/backwards links from <em> to </em> like in a skip list.
I have implemented this list, but never actually used it that way. It became too confusing. So I now remove mismatched tags, while creating the list, making it equivalent to a tree. Now I probably should rewrite it to not actually store closing tags.
The thing I dislike most is the treatment of self-closing tags like <br />. In HTML5, the correct form is <br> but you can add the slash and it is ignored as syntactic sugar.
But you cannot write <div id="mydiv" />. In XHTML, this was a shorthand for <div id="mydiv"></div>. It's really ugly that you have to write the closing tag in this case, and for seemingly no reason.
https://stackoverflow.com/questions/3558119/are-non-void-sel...
Sure you can. And exactly as with <br/> the / is ignored.
Ignored means exactly that. HTML does not have xml-style self-closing tags at all.
> for seemingly no reason.
The reason is pretty simple: there were no self/closing tags in SGML, so there were no self-closing tags in the mangled mess browsers made of it.
The goal of the HTML5 parser effort was to create an interoperable and backwards compatible parsing algorithm. Adding self-closing tags to the langage would have tin counter to that goal and would definitely have broken existing code unnecessarily, even though it might have been convenient in some cases.
> Sure you can. And exactly as with <br/> the / is ignored.
Obviously, the phrase "you cannot write" has the context "as a shorthand for self-closing the tag", and doesn't deserve some pedantic "of course you can do that, it just doesn't do what you hoped it would" response :/.
To really be pedantic about it: SGML actually had a feature similar to this feature and HTML specified it; however, the syntax "<br/" was a valid way to write that tag (technically an open tag, but due to the DTD it would auto close) and the ">" after would get rendered... not ignored (and in fact might even cause the document to fail if that > was in a disallowed context of a strict document).
https://jkorpela.fi/html/empty.html
Ironically, the W3's comment on this is to discourage its usage "until widely deployed", as they have essentially always specified stuff no one implemented correctly (lol).
If and when browsers drop their XML parser, I hope they will just parse these docs as HTML and only then would I loose the option of self-closing tags.
You are answering my complaint by repeating the fact that I complained about? :-D
> The goal of the HTML5 parser effort was to create an interoperable and backwards compatible parsing algorithm. Adding self-closing tags to the langage would have tin counter to that goal and would definitely have broken existing code unnecessarily, even though it might have been convenient in some cases.
I'd wager that in pre-HTML5 code, the overwhelming amount of cases of <div /> were supposed to mean <div></div> and not <div>.... Why would you even put the slash there if you wanted an opening tag? HTML5 would have been a great opportunity to legalize the former behavior, and that would have broken less code. I distinctivly remember that the self-closing worked as expected even on non-void-tags in quirks mode (not XTML) in the mid 2000s, because at some point I had to fix them after new browser versions came out.
The problem with being lax in what you accept is that it builds immense complexity because you have to deal with a potentially infinite number of special cases.
But this books appears excellent. I read the encoding chapter just now, and it was very engaging and informative, explaining the historical reasons why things are the way they are.
For example, it says that meta tag about character encoding were initially meant to be read by the server so that it could set the proper encoding header... Which they never did... I did not know that!
It seems no web server ever implemented it; it would probably have been a little beneficial because each file would only need to be pre-parsed when it is written on disk (or first read by the server), instead of this being done by every client on every load.
It makes me think, maybe other kinds of analysis and normalization could be done by the server, which would allow for simpler (and stricter) browsers.
Also, you can still write HTML5 as XML (XHTML5). It's not talked about much, but it's totally a thing.
Just what I was looking for.
> The following are equivalent: <link rel="stylesheet" href="style.css" /> [and] <link rel="stylesheet" href="style.css">>
No they're not. The author should study the 1998 WebSGML adaptations to ISO 8879 (the SGML spec), and in particular the NETENABL IMMEDNET feature.
I mean, I welcome discussion about HTML parsing; just not the kind of narrative painting HTML5 as some kind of academic or community consensus when it's very clear that HTML5 as it is is a dead end, and a regression due to its monolithic nature tying a parsing spec to a concrete markup language, unlike SGML and XML.
Edit: I read this but it didn't add much clarity -- https://medium.com/hackernoon/to-close-or-not-to-close-4365d...
Loads of changes may have happened to CSS and JS, but HTML has remained largely the same, save for the addition of a couple elements (section, article, main, header/footer added by Ian Hickson early on without any input by other parties) and import of SVG (the major contribution of XML) directly into HTML.
Which is exactly the problem: HTML hasn't advanced for eg mobile, but everything else around it has turned into absurd complexity so we can continue to pretend we're writing 1990's casual academic content (the original use case for HTML the markup language). When, of all the languages in the web stack, markup is the one thing that has typing, schemas, and a schema evolution story.
I agree no one should base a new markup on HTML5, but eh, no one is? You seem to be attacking a strawman there.
Mind that the author is literally one of the designers and maintainers of the HTML5 parsing specs. From the section "About the author":
> He contributed to the design of some aspects of the HTML parser specification […] and is currently an editor of the WHATWG HTML standard and the WHATWG Quirks Mode standard.
WHATWG (and W3C) isn't exactly a shining example of succesfull standardization when measured in terms of broad community involvement given their pay-as-you-go statue, or measured by adoption considering there are only about two browser code bases left. It has been "whatever Chrome does" for years now, not to speak of being a monolithic spec the size of the NYC phone book and a process having never achieved any final published version after 17 years of work.
Edit: Also, while WHATWG may not be a resounding success, WHATWG's HTML5 parsing definitely is. We now have large number of interoperable implementations that can parse existing HTML documents in the wild. It is one of the most dramatic success of standardization I have witnessed.
(b) The reason you fork a codebase is to make it different. They've been forked for eight years or so now. This is longer than many codebases even exist for! I don't know how different they are in practice (especially in the matter of parsing), but I think it's unreasonable to assert that just because they're forked, ipso facto they are not independent. Instead, you should have to show that they are not independent and then explain this in terms of common origin.
edit: To add to this point (b), afaik the old Netscape and IE had a common origin, in that they both derived from NCSA Mosaic. But it would be ludicrous to imply that they were in any sense equivalent!
By forking, you inherit a code and test base encoding all quirks and warts (= that which is actually spec-worthy), but don't gain any insight into the spec quality or alignment with a spec at all.
Now, of course, we have a much worse problem with web specs: that no browser has been started from scratch since basically forever. Thus, the spec can serve merely as an after-the-fact documentation of existing browser behaviors, or at best record an intent for "browser vendors" to implement some feature. Which is what WHATWG basically is and how it came to be: a platform for "browser vendors" to agree on evolving web standards without the ceremony of a full-blown spec process like W3C's. There's nothing wrong with that; it's just not what should be called "a standard", especially if it doesn't produce a published, versioned document or other deliverable but is always work-in-progress by design.
Tmk, the death of Trident/EdgeHTML is a little overstated, since it is still used by the UWP embedded browser widgets. This is probably more of a concern for some subset of js/CSS library developers than for app developers.
Very different in some aspects. Not that different in others. Depending on where you look.
They have different Javascript engines and security models. But HTML and CSS is mostly the same/similar.
But then you have Chrome implementing over a 1000 more Web APIs [1] some/many of which will never see the light of day in other browsers [2]
(The argument regarding the dominance of Google has merit, but these are the realities of modern web politics. Let's see, whether "portal" will be accepted or not, which will be a valid test case.)
Have a nice day folks!
Unless you have an ambiguous system which:
- has to produce valid output in the presence of invalid input
- apparently has to take external scripting such as `document.write` and `document.innerHtml` into account
- has decades of legacy and conflicting standards thrown in
And, undoubtedly, it's still the easiest part of a browser.
HTML5 did bring more order to the chaos, but I doubt it reduced it that much.
In the past you could write a straightforward parser for the pages you had.
When it did not work on some other pages, you could say, that page has invalid HTML, it is not expected to be parseable.
Now you have to spend a lot of time to implement everything in HTML5
(Actually, this isn’t true in the presence of scripting because of document.write(), which allows you to feed arbitrary data to the parser during parsing. But that’s mostly a technicality.)