The evolution of the web, and a eulogy for XHTML2
devever.net
devever.net
Indeed, I am consistently surprised at the document-display features that are still missing from web browsers that are increasingly preoccupied with serving as application runtimes.
The web was originally envisioned as a platform for scientists to share documents, but thanks to Google dropping MathML, there's still no cross-browser, non-kludge way to display math.
Browsers still can't justify text properly. Even with hyphenation (which iirc is still a problem in Chrome), the greedy algorithm used for splitting text across lines still results in too much space between words when justified.
There seems to be a complete lack of interest in Paged Media support, so if you want your web-page to be printable with nice formatting, you basically need to provide it as a PDF.
I've gotten to the point where I'm often happier to read something in-browser as a PDF than as a web-page. Sure, it can't reflow properly, but at least it won't have a 2-inch sticky header and make XHR requests to 10 different domains. My vision-impaired father reports that his text-to-speech software now works better with PDFs than with real web-pages. This is insanity.
Chrome and only Chrome has an include extension, but for security reasons (something about cross-site foo), you can't use it when testing an HTML file on your disk, because you need a hostname to check the same origin policy. You need to put it into a local webserver, which, fortunately, is fairly easy with Python.
In the early days, when dynamic webpages were processed by something called CGI and even when we started using PHP, this was like 90% of the use case for doing a dynamic webpage. The last 10% might have been some sort of counter.
> I never understood why there wasn't an <include>
> element in HTML, that would just paste the content of
> the linked document. Why wasn't such an element added
> in the late 90's or early 2000?
As far as I know, for SGML (and therefore old fashioned XML, too), 'include' gets realized via parsing of external &entities;. That's how the Mozilla XUL applications (Seamonkey, Firefox, Thunderbird) etc. realized localization, for example.I don't know, when XInclude has first been ratified (the current, 2nd revision, of the spec is from 2006), but with it, a generic inclusion mechanism exists for XML, therefore also for XHTML.
With XHTML you could (in theory, as said, the spec is there) use XInclude.
<?xml version='1.0'?>
<html xmlns="http:www.w3.org/1999/xhtml"
xmlns:xi="http://www.w3.org/2001/XInclude">
<head>
</head>
<body>
<p>120 MHz is adequate for an average home user.</p>
<xi:include href="disclaimer.xml"/>
</body>
</html>
I am not quite sure, yet, how namespace mixins are handled in XHTML, I think, one must define a new schema (which, honestly is a crappy requirement and should be abandoned), effectively creating an XHTML+XInclude document type, but that may be only a real issue for strict validation.In my purely HTML github pages I ended up hacking a header into the common CSS file bc that's "included" everywhere.. it's ugly and I don't think you can even put links in it... but it sorta works. If I wanna tweak it then it's all in one file
PS: this is a naiive question from a non-web dev
I think you can disable that check as a command line option. I cannot check it right now since I don't have chrome installed however I think it was --allow-file-access-from-files .
Granted, I'm surprised XSLT is supported at all in modern browsers, and who knows how long it will continue to be supported, but it does work now for the basic "menubar at top" use case.
I'm extremely excited by Igalia's work on MathML and glad that the Google Chrome team had given it their preliminary approval, but in the context of criticising Google for working on the "application web" and ignoring the "document web" (for scientists and others), it doesn't change anything[0]. Google is neither doing the work, nor funding it[1], and their contribution until now has mostly (solely?) been to be open about merging it.
[0] not that you explicitly said that it did...
[1] It's funded by the NISO and the Alfred P. Sloan Foundation.
Second, because you want to print -- in which case, your artifact ends up in a page.
Third, because a screenful, even if dynamic (based on resolution, etc) is still a kind of a "page" (in HTML parlance, a "viewport"), and you might want to design to take advantage of what fits in one.
I've tried using it for my resume since may about 10 years ago, to make it possible to convert it to PDF with a standard browser, but to this day many of these features aren't fully supported.
That's kinda the main selling point of software like Prince, which has implemented all of this ages ago, but it requires a licence.
> There seems to be a complete lack of interest in Paged Media support, so if you want your web-page to be printable with nice formatting, you basically need to provide it as a PDF.
These are the two main reasons I think of whenever someone complains that scientific publication should be Web-based instead of PDFs. Browsers are still generally lacking when it comes to proper typesetting.
On the gripping hand, in a century those reconstituted dead trees will still be legible by anyone who can lay hands on them; the likelihood that an interactive presentation from 2019 will be legible to anyone other than a computer archæologist in 2119 is effectively nil.
Heck, there’s a decent chance that it won’t work in a year or two!
I have enough experience with PDFs and screen readers that I find this very surprising. Which PDF reader does he use, and does he get most or all of his PDFs from a particular source?
You might try:
word-spacing: -0.1ex;
Good justification of text requires that sometimes the spacing between words is a little narrower than normal. Perhaps browsers are unwilling to do this unless you explicitly give them permission. But I wish they could just adopt a good text-flow algorithm like Adobe InDesign had.But good justification also requires hyphenation. And last I looked, browser support was spotty. It did not work on Chromebooks, for example.
Probably a huge tangent, but UnicodeMath [1][2] is at least encodable and (basically) readable just about everywhere UTF-8 (et al) is supported, which is all browsers today, if unfortunately only just a linear presentation by default. I'm surprised there aren't more progressive renderers for it to up-level it to more traditional two-dimensional renders in browsers directly yet (after several years being standardized as a Unicode Technical Note). A cursory glance shows that even MathJax still doesn't seem to support UnicodeMath (at least out of the box), like it does MathML and AsciiMath. (But I think raw UnicodeMath is easier to read than MathML or AsciiMath in a "progressive" fashion from linear to "professional"/two-dimensional rendering. Though that's a personal preference/aesthetic that can get quite subjective.)
[1] http://www.unicode.org/notes/tn28/UTN28-PlainTextMath-v3.1.p...
[2] https://blogs.msdn.microsoft.com/murrays/2016/09/07/unicodem...
> If anyone constructed a PDF, which was itself blank but, via embedded JavaScript, loaded parts of itself from a remote server, people would rightly balk and wonder what on earth the creator of this PDF was thinking — yet this is precisely the design of many “websites”. To put it simply, websites and webapps are not the same thing, nor should they be. Yet the conflation of a platform for hypertext and a platform for applications has confused thinking, and led developers with prodigious aptitude for JavaScript to mistakenly see mere websites of text as a like nail to their applications hammer.
Now, maybe they never should have allowed Flash and friends, but the genie was well out of the bottle. You could either have HTML and JS based functionality, or the plugins.
Instead they infected the HTML standard with everything Flash was being used for.
The OA states in his article, that he is fine with many of the new APIs. It's not HTML5 vs. XHTML, it's 'applified web' vs. 'document web'. And I agree.
We could get some of the advantages if we just started shunning JS and telling people ("Proudly Javascript-free and free of third party trackers! You might have noticed that this site loaded quicker than other sites you have seen lately. Mostly this is because we don't bug your machine with insert-carefully-crafted-explanation-here.")
Unfortunately, this is also the case for many PDFs amongst some of the more niche use-cases. Like NDA-bound many-thousand paged specifications. The PDF standard lets you do it, so some do. Not all readers actually support embedded JS, but most do in some limited capacity.
In the case of a physical printer, it's usually set up so there's [redacted] instead of blankness (that the JS replaces), so you get redacted images and paragraphs.
In the case of software printing to a static PDF which you can then print, the JS payload is generally set up to embed a series of rather obvious markers so that if the static copy is ever publicly revealed, they can go after you (DRM is easy when you only have a dozen customers in the entire world), but most PDF software that supports JS also supports JS disabling printing features so that little disincentive is unnecessary.
https://www.owasp.org/index.php/XML_External_Entity_(XXE)_Pr...
> rather than articulating particular requirements and principles but not how they need be met, the WHATWG specifications tend to be written in a highly algorithmic and prescriptive style; they read like a web browser's source, if web browsers were written in natural language.
It turns out that if you want to have pages work the same in every browser you need to have every browser doing the same thing when it interprets the pages.
> The pursuit of the semantic web has changed in the era of HTML5, which represented a rejection of XHTML — to me, a seemingly bizarre rejection of having to write well-formed XML as somehow being unreasonably burdensome.
In practice, people won't write valid XML. We had a lot of cargo-culting, people putting self-closing tags into HTML, but people weren't using XML editors. And without an editor that understands XML it definitely is unreasonably burdensome to create XML. What we saw instead was that even most "XHTML" documents were not valid XHTML and were served with an HTML content type. If you had served them instead with an XHTML content type the browser would have simply refused to render them.
Both of these were recognitions that the previous approach wasn't working, and that if the spec was to achieve its goals we needed to try something different. Under WHATWG the spec has moved from "yes, the spec says this but it doesn't matter" to "the spec describes what the browsers do, and the browsers treat cases where they violate the spec as bugs". Sites now really do work the same across browsers, and WHATWG deserves a lot of credit for that.
Yes, but the change was that first browsers adapted to the standard, now the standard adapts to the browsers. That's the big change.
On the other hand, an independent standard organization (which, of course, the w3c isn't) would not have to follow a browser's vendor agenda. They'll ideally want to do what's right for the web.
To sum it up, I don't think having browser vendors set the standard is in the best interest of the web platform, especially when they are so unequally represented in actual usage (and we can debate about how we got ourselves in this situation).
Then: if I wrote to the spec I would find the major browsers would all handle my page differently. Writing complex cross-browser pages was a constant pain involving IE-specific hacks (special CSS comments that only IE understood). The browsers were not interested in implementing the spec because it meant a ton of work for no benefit and existing pages would break, and the W3C had moved on to XHTML-only approaches.
Now: I can write to the spec and the major browsers (based on WebKit, Blink, EdgeHTML, and Gecko) will all do the same thing with my page. Spec violations are bugs, and are taken seriously by the browser vendors.
(Disclosure: I work for Google, though not on Chrome)
Disagree. You open the tag, you close it. Nested. I worked with a non-developer who had some KML (https://developers.google.com/kml/) dumped on her without any training whatsoever and when I explained a few basic thing, including the above rule, she got the principle of it in a few minutes. Because it's not hard (although dumping a KML task on a non-dev is a bastard thing to do).
> but people weren't using XML editors
I personally think that lowering the bar to let as many people in as possible isn't necessarily a good idea. Just IMO
I personally agree with you, and still close my tags, use quotes around attribute values, and provide empty/duplicate values when needed.
I also still use XML for data storage, and used ColdFusion back in the 6/7 days, so that probably had a influence on my choices.
That's probably the most popular argument against X(HT)ML and in favour of HTML5. In fact I think among the popular programming/markup languages, only very few are so forgiving. Namely it is HTML(5), JS, CSS and perhaps Shell script and Perl. But even in these cases following best-practices and using Linters has become extremely popular. On the other hand you have strongly typed languages or even languages like Python or Makefiles that even make sure you use consistent whitespaces.
I think nearly everybody uses quite powerful editors with a load of plugins these days because accelerate editing and also do autoformatting.
> Sites now really do work the same across browsers, and WHATWG deserves a lot of credit for that.
On the other hand there are just 2 popular/"usable" browser engines left. I think XHTML is far more modular, maybe it would even be possible to outsource some browser rendering tasks to XSLT transformations. HTML at one point became a messy standard through the Browser competition and then WHATHG somehow manifested that situation I guess. Now there's a massive monoculture of Browser engines.
I count at least three
With that said, I would love the author to go on and elaborate on other advantages of XHTML2, such as possible integrations with XForms (including more inputs and sending requests without page reloading and without JavaScript), XFrames, the single header element <h>, every element as a hyperlink, etc. Then there are MathML and XSLT. If XHTML2 became a reality, we would probably see XSLT 2.0 more actively adopted by the browser vendors, which is a good thing in my book.
That's not the fault of the XML community at large, but just a lack of resources for the implementation of an unpaid open source project.
You can happily use XPath 3.1 with Saxon, BaseX, eXist. All three use Java, so, it's not portable, but Saxon has a C library, that mirrors the Java version 1:1, and that C library is also available as open source, though, it lacks some XPath 3.x features, like higher order functions, then.
For the command line, there is a partial XPath 3.1 implementation with 'xidel'.
But I agree, libxml and libxslt being at XPath 1.0 for so long, did not serve XML well.
The fragmentation you mentioned is part of what made this so frustrating: if everything you used was within certain toolchains, the experience was fairly good but then you'd need to use a different ecosystem and either drop back to good old XPath 1 or take on more technical debt. In many cases, the answer I saw people favor was leaving the XML world as quickly as possible, which is something the community has a strong interest in.
1. For example, Saxon added support for Python just a few days ago: https://www.saxonica.com/saxon-c/release-notes.xml Imagine if that had happened a decade ago and everyone who was stuck with libxml2 could have easily switched?
On a fair note, one should also take into account, that Saxonica is a rather small, even if highly skilled, shop and the program, they create, is a huge undertaking.
Handling namespaces correctly requires that the parser API be changed in non-backwards-compatible ways. This would have broken every single piece of code that used an XML parser. So instead, people mangled documents by simply flattening all the namespaces together if you tried using the old API.
This was a godawful nightmare and made everybody who was around at the time absolutely hate XML namespaces -- even those of us who know why they are so important. Plus namespaces are not exactly an "ELI5" topic, so a lot of lazy programmers looked at this and said "that's complicated, I don't want to learn it, HEY LOOK there's this older deprecated API that doesn't have them -- I'll use that!" So the old APIs became immortal and in fact gained additional users long after they were deprecated.
They should never have let the 1.0 standard out the door without namespaces in it.
1. User gets a simple XML file and writes XPath, XSLT, or other code which says `/foo/bar`, which fails.
2. User notices that while it's written as `<foo><bar>` in the source, it's namespaced globally so they change code to use `/ns:foo/ns:bar`, which also fails.
3. User does more reading and realizes it needs to be `/{http://pointless/repetition}foo/{http://pointless/repetition... or something like repeating the document-level namespace definitions on every call so their `ns:foo` is actually translated rather than treated as some random new declaration.
4. User does something hacky with regular expression to get the job done and at the next chance ports everything to JSON instead, seeing 1+ orders of magnitude better performance and code size reductions even though it's technically less correct.
That experience would have been much less frustrating if you could rely on tools implementing the default namespace or being smart enough to allow you to use the same abbreviations present in the document so `<myns:foo>` could be referenced everywhere you cared about it as `myns:foo` with the computer doing the lookup rather than forcing the developer to do it manually.
JSON is a better data language, but a lot of that can be laid at the feet of XSD and the big squandering of momentum it represented.
It was a page that displayed IRC logs, with linkable anchors for each line, automatically breaking words that were too long for the browser, turning text into links by regex, interpreting terminal colors etc.
Each time I had a tiny cross-site scripting bug in there (some part wasn't XML-escaped properly), some data would eventually trigger it (you wouldn't believe the amount of encoding junk on IRC), the browser would simply refuse to render anything at all. Inconvenient for my users, but it made sure such things didn't slip by unnoticed.
---
This was a side project, done for fun, and a few friends that used my site. If I had been trying to make money from it, the first thing I would've done is to switch to something less strict, so that tiny errors wouldn't stop rendering the whole page.
It was very nice to load an XML document and look for tags in a specific namespace instead of using a specialized HTML templating engine.
> The Extensible Markup Language (XML) is a subset of SGML that is completely described in this document. Its goal is to enable generic SGML to be served, received, and processed on the Web in the way that is now possible with HTML. XML has been designed for ease of implementation and for interoperability with both SGML and HTML.
The "generic" part refers to XML being canonical, fully-tagged markup not requiring vocabulary-specific markup declarations for tag omission/inference, empty elements and enumerated attributes like is necessary for HTML and other SGML vocabularies making use of these features.
That XML has failed on the web doesn't mean one has to give up structured documents. In fact, HTML can be converted easily into XHTML using SGML [1]. If anything, markup geeks should embrace SGML (an ISO standard no less) to discover the power of a true text authoring format. For example, SGML supports Wiki syntaxes (short references) such as markdown.
[1]: http://sgmljs.net/docs/parsing-html-tutorial/parsing-html-tu...
Look at this "<p<a href="/">first part of the text</> second part". This is a valid document fragment in HTML 4.01 because HTML is authored in SGML.
Writing a correct XML parser is much easier than writing a correct SGML parser, and what's more important, it's much easier to recognize errors.
I agree with OP that HTML5 should have been XML from the start. Nowadays, you hardly write any HTML by hand and even if you do, it's easy to write syntactically correct XML.
It's true that you can convert any HTML into XML with ease but it's still a stupid, unnecessary step.
The point is that this hasn't happened; neither back in XML's heyday, and much less today. Now you can bemoan XML's demise until the end of time, or you can fallback to XML's big sister SGML. As I said, SGML has lots of features over XML that are in fact desirable for an authoring format, such as Wiki syntaxes, type-safe/injection-free templating, stylesheets, etc. on top of being able to parse HTML. Many of these features are being reinvented in modern file-based CMSs and static site generators, so there's definitely a use case for this. Whereas editing XML (a delivery rather then authoring format) by hand is quite cumbersome, verbose and redundant, yet still doesn't help at all in how text content is actually created on the web.
XML, while not the main format for modern HTML, isn't dead, so why would one bemoan it's demise? In fact, it's quite widely used.
SGML is needlessly complex as an authoring format. Even HTML was considered too complex and that's why we got lightweight markup languages like MarkDown and AsciiDoc.
I would be very surprised if we ever turn back to something like SGML. Especially as there are well designed LML as AsciiDoc or reStructuredText.
[1]: http://sgmljs.net/docs/producing-html-tutorial/producing-htm...
The key requirement for HTML5, And why it succeeded where XHTML had limited success, was that existing HTML docs had to work with it. Which is why it has both an HTML and an XML format.
It was not wrong for it not to be pure XML, it was absolutely necessary.
XHTML2 OTOH was a convoluted mess of modular documents that couldn't get out of design hell and completely disregarded the developers of browsers.
HTML5 succeeded because it was implemented by the browsers, because the same people that programmed the browsers were on board when designing HTML5.
You could and a lot of people _tried_, or at least pretended to. But the vast majority of documents that tried to do this failed to actually be well-formed XML, for various reasons... In practice, even restricting parsing as XML to cases when the page was explicitly sent with the application/xhtml+xml MIME type would leave a browser with problems when sites sent non-well-formed XML with that MIME type. This was a pretty serious problem for Gecko back in the day when we attempted to push XHTML usage (e.g. by putting "application/xhtml+xml" ahead of "text/html" in the Accept header). So we stopped pushing that, since it was actively harming our users...
SGML is the only game in town able to parse (a significant part of) HTML based on an actual standard, and is also the only realistic perspective for folks interested in the web as a standardized communication medium going forward.
HTML5 recognises that there was this gulf between the specification and the actual usage and sided with real world usage.
Xml is also structurally so much simpler than xml. Making an editor able to auto indent xml is simpler than making it auto indent html.
HTML5 just standardized all these quirks, leading to a uniform parsing model instead of an even bigger x-browser mess.
As far as XML “tools”, I am shocked at how even now I encounter real XML parsers that don’t necessarily reject malformed data files but do atrocious things with them (like silently pretend that certain tags were not even in the file). Thus, I end up using extra steps like a linter as a front-end sanity check. And while this example is a pure-data application, a linter is also a sensible front-end sanity check for HTML. XML isn’t going to win over HTML if it requires the same steps to clean up imperfections in the process.
A lot of people used to emit XHTML with invalid string builders rather than XML serialisers though. Much XHTML was ruinously broken.
This is a classic "worse is better" situation. HTML5 may be "worse" than XHTML, from the standpoint of extensibility, namespacing, code cleanliness, and so on. But HTML5 is simpler to write for people who knew HTML4, and easier to get right using the one ubiquitous web development practice: staring at the rendered result in your browser, which every web developer has installed. So it's "better", and ends up winning.
Had the XHTML2 standard been adopted by many, browsers would have still had to support all the other non-X HTML documents, which would have never disappeared.
HTML5, with the exception of the new elements, just formalized the existing web parsing strategies for better cross compatibility.
As a web user, I don’t miss the days of XHTML sites randomly completely breaking because some tag wasn’t closed.
As a developer, I only miss XHTML2’s support for `href` on any element.
Ironically that was a very rare event in practice. For it to even possibly happen, three things had to come together which were still uncommon even at XHTML's height:
1. The page had to be written as XHTML
2. The page had to be served as XHTML
3. The page had to be parsed as XHTML
Usually at least two of those things weren't happening.
The vast majority of sites were "HTML 4.01" others were at best "XHTML 1.0 Transitional" (which in practice meant the same thing). Those using pure XHTML were relatively few. And of those who did, no major site served it as such because it would have locked out IE users, IIRC.
This! I would go a step further, even. The "web-applifier" community should just leave the classic web and do their own:
* protocol (I am sure, HTTP is not ideal for serving apps) * GUI description language (document markup language for UI design, really?) * runtime (let them have WebAssembly and whatever they need) * each app then could have their own window, making it look like a traditional app
The reason HTML5 became what it is is because many people wanted to see the open web thrive as a competitor to closed, controlled eco-systems like the mobile application development platforms.
Just 10 years ago this was a mainstream view - there were groups who were fighting to give browsers web cam access so they could be used as video-messaging platforms, groups fighting to give location access so we could write location aware documents and apps etc etc.
I still believe this was the right decision.
So in that context I do agree with the article.
One of the reasond a platform succeeds is because it can be many things to different people.
I completely rejected the idea that there should be separated document and application models. Instead I think there are capabilities which should be layered to add abilities.
This seems like a lack of imagination. The modern scientific publication Distill.pub makes heavy use of AJAX (eg https://distill.pub/2019/activation-atlas/) and it's easy to imagine it using a webcam (eg, to demonstrate semantic segmentation)
Now I hope that the form extensions proposal [1] will gain some traction.
[1] http://cameronjones.github.io/form-http-extensions/index.htm...
the whole idea of a web of semantic hypertext was never a reality. that idea died with gopher. (anyone remember that?)
why? because gopher had a builtin navigation system that allowed you to manage directories and document structures that didn't belong to the documents themselves. semantic hypertext within gopher would have worked well. as would have applications. but without gopher we were forced to reinvent that navigation and squeeze it into our documents, overloading them with stuff that didn't belong inside.
think of a library with books. what the semantic hypertext promised was to make all those books into interactive texts where you can easily jump from one reference to another. but those books still need a library to live in.
what the web ended up doing was to remove the library completely, forcing me to reinvent the library within the book. suddenly the semantic document that i want to send you not only contains references relevant to its context, but it has to include the whole navigation for my library, because there is no way to do that externally. with that navigation included, you are no longer getting a semantic hypertext, but an application.
now, we get to write that application in javascript and actually run it on your device, instead faking it on the server. but on the flip side, on the server i can now finally go back to serving static documents. i can finally serve semantic hypertext documents as they were meant to be served because i can separate the application from the content, and i can treat the content as static as it was meant to be treated.
i am not reinventing navigation logic in javascript. browsers never had navigation logic in the first place. gopher had that. i always had to re-invent navigation logic for every site i built and was forced to embed that into html dynamically so that site visitors could find their way.
WHATWG did not "usurp" the W3C. The W3C abandoned HTML development by focusing on XHTML2, incompatible with HTML. This left an opening for someone to propose backwards-compatible extensions of HTML, since the W3C was explicitly not interested in that. WHATWG was formed to do this and produced HTML5. HTML5 was adopted by industry, XHTML2 was not. In an attempt to stay relevant, the W3C tried to stage a hostile takeover of HTML5. That attempt failed because the W3C had blown their credibility by that point. However there were still some advantages to having a W3C-approved HTML spec, so an agreement was reached where the W3C could approve the specs produced by WHATWG.
There are technical reasons why HTML <object> was not suitable for audio and video elements. For example, media elements need to expose media-specific JS APIs (e.g. seek()), but the MIME type of an <object> can change over time due to URL loading and DOM attribute changes, which would mean that the interface exposed by the element would need to change unpredictably over time, which would be a nightmare for developers. Also, there were very nasty legacy browser compatibility constraints around (mis)use of <object>.
The author misunderstands, or misrepresents, the WHATWG's spec design philosophy. Unlike the W3C, the WHATWG treated compatibility with existing Web content as essential. That means the WHATWG specifies existing browser behaviour where there is significant existing Web content that requires it. The W3C, on the other hand, tended to assume that Web developers pay attention to specs and that writing down conformance requirements would magically cause all Web content to be updated to satisfy them. Those assumptions are not true. (The idea that Web developers would migrate to XHTML2 because the W3C proclaimed it as the future was in the same vein.)
XML syntax for HTML failed for various reasons but not because of the WHATWG or browsers, which always supported XML syntax for HTML. One major problem is ensuring that dynamically generated XML pages are always valid XML. It is very easy to have bugs so that under some conditions (e.g. malicious user input) the server outputs invalid XML and produces a "yellow screen of death". Common examples of those bugs were bugs that allowed the Unicode 0xFFFE or 0xFFFF code points to slip into the XML output, which are not allowed in valid XML. A similar problem is when users interrupt a partial download of an XHTML file; the file has unclosed tags, so a conforming browser will replace the partially loaded and rendered document with a "yellow screen of death". This is not what users or developers actually want. (This assumes the browser bends the rules to allow partial rendering in completely loaded and validated XHTML documents, which is something users and developers do actually want.)
HTML 5 started out as an W3C position paper to start the work of creating a backwards compatible successor to HTML 4 backed by several browser developers. The proposal for this was rejected in favor of continuing work on the non backwards compatible XHTML 2. As these browser developers had a need for a spec, the WHATWG was formed to create what became HTML 5.
Several years later when it became clear that their specification was the closest thing to describing what browsers actually do, W3C basically endorsed it as a recommendation. However, the WHATWG continues to drive work as it has been highly successful in producing the high quality specifications required to achieve the high levels of interoperability between the remaining browser engines.
XHTML 2 gradually became redundant as most of the functionality that web site developers actually needed from browsers got absorbed by HTML 5. As browsers could implement spec changes as they were happening, within a few years of the working group forming, it had become a huge success in standardizing many new features.
At the same time, many mobile browsers simply disappeared as Webkit (a fork of KTML by Apple) and the Chrome fork of Webkit by Google became the norm on mobile. This was important because many W3C working groups were dominated by people from the mobile phone and telecom industry seeking to control the standard for long forgotten things like WAP and the various mobile profiles of XHTML. Once decent browsers appeared on mobile (i.e. browsers that supported HTML 5), the need to continue to do work on XHTML 2 disappeared. Most of the companies that then dominated the mobile web were out competed by Google and Apple; neither of whom was a major mobile player at the time the WhatWG was formed. By the time W3C endorsed HTML 5, Android and IOS were dominating the mobile web to the point that even MS threw in the towel by first creating Edge (an html 5 browser with no backwards compatibility for IE specific stuff), and then recently just switching to Chrome entirely.
I also don't, and never really cared much for "semantic hypertext," linked data, XHTML, XHTML2, RDF, nor is my browser usage predominantly about just sharing/receiving information.
That being said, there's so much involved in writing a Web browser that writing one from scratch would take years. There are only a few active implementations, and I don't expect that change for a long time, if ever (at least not without a lot of funding). So I can understand the author's viewpoint.
What do you use the browser for then? Do you primarily play games/music/movies?
You mention applications in your comment; if that's not sharing/receiving information, what are you doing in them?
I should have been a bit more clear about "sharing/receiving information," or rather, reworded it as "viewing static documents."
In fact, I'd go one further: If gopher had taken up multimedia quicker and then beat out the web, there would be no Google today.
These sort of things are not inherent to the web. People and entities want to transmit data in the format(s) that are easiest for them. It's up to the aggregators to make sense of it all.
It requires some investment to come up with tools for analyzing files. I’m wondering if the tools would have been as sophisticated if documents had lowered their barriers. For example, how do you justify developing a machine-learning model to look for more, if it seems all your documents are already semantically tagged with the details that were important to somebody?
https://github.com/mozilla/hubs-cloud/wiki/The-Web-Emergent-...
And many times that choice is made for you: there’s a link to a Vulture.com post on HN right now (about Disney archiving the Fox catalogue) and every time I try to read it, Chrome crashes trying to keep up with the volumes of JavaScript they add, ostensibly to keep advertisers happy.
https://webpagetest.org/result/191027_BF_57a6cd57fc6fe629628...
https://webpagetest.org/result/191027_2G_fba74afe0c488b99e53...