1. <a b="42>c">d
2. <a/b/c=d/e>f
3. <a/="42>b
4. <a x=&0>&0</a>
5. a<!--->b<!--+->c<!-->d
Really, don't try to answer and just use complaint HTML parsers. 1. <a b="42>c">d
2. <a/b/c=d/e>f
3. <a/="42>b
4. <a x=&0>&0</a>
5. a<!--->b<!--+->c<!-->d
Really, don't try to answer and just use complaint HTML parsers.I'm really sad that they didn't go with a XML base for HTML5.
I'm really sad that they didn't evangelise an XML base for HTML5, and that many HTML5-ish tools don't explicitly support XML, but it's not strictly true that they didn't go for an XML base for HTML5[0][1]
That you can put the same information of a HTML5 document into a XML document doesn't help much if most of the HTML5 documents out there are not polygot.
> When a document is transmitted with an XML MIME type, such as application/xhtml+xml, then it is treated as an XML document by web browsers, to be parsed by an XML processor. Authors are reminded that the processing for XML and HTML differs; in particular, even minor syntax errors will prevent a document labeled as XML from being rendered fully, whereas they would be ignored in the HTML syntax.
HTML does use an XML base (elements, attributes, namespaces, &c.), it just doesn’t use an XML parser most of the time. But the XMLness is easily observed in various DOM APIs.
I intended the latter. In fact I'm a bit surprised that I have ever been asked for this, I thought "HTML" nowadays exclusively refers to HTML5...
But why should I? Who writes HTML by hand these days?
The SGML heritage of HTML 4.01 and earlier lead to some gruesome legal constructs that look surprisingly similar to your examples. Looks like every generation has to make their own mistakes.