I suppose I should actually give more details. You probably already know most of this history.
I was basically still a kid, working for a vendor in 1997, and my job was to write a library for manipulating SGML parse trees. Today, it would be a DOM library, but the HTML DOM wasn't standardized until the fall of 1998. I think I remember hearing about HyTime, but I didn't have a copy of any of their related work. I basically had a copy the SGML standard, and of James Clark's nsgmls parser. A very experienced and capable colleague was in charge of writing an SGML parser. Our target machines had roughly 8MB of RAM and maybe 80 MHz CPUs. Which meant they could parse SGML, but it took a noticeable amount of their resources. My frustration with implementing SGML was first-hand, although I was working on the easy bits.
SGML at the time was mostly an expensive enterprise niche, except for James Clark's excellent free tools. The most popular application I saw in the wild was probably DocBook? A lot of SGML vendors were playing up the connection between SGML and HTML in an effort to seem relevant, but their tools were notoriously expensive. None of the vendors were especially mainstream.
The arrival of XML really was a sea change. Judging from my colleagues who actually attended trade shows and who paid attention to the other players, the enterprise SGML vendors basically saw XML as an attempt to break out of their niche.
XML was almost immediately "mainstream", at least for people who wanted something other than HTML tag soup. Of course, as you point out, HTML ultimately turned away from XML. But the number of people who knew and cared about XML seemed to rapidly exceed the size of the SGML community.
> Finally, SGML is the only game in town able to parse HTML based on an international standard.
Is this actually true? Can any standards-compliant SGML parser actually parse arbitrary valid HTML 5 as SGML, and pass the test suites? https://www.w3.org/html/wg/wiki/Testing
The HTML 5 spec at https://html.spec.whatwg.org/ states:
> Also, since neither of the two authoring formats defined in this specification are applications of SGML, a validating SGML system cannot constitute a conformance checker either.
HTML 5 has two syntaxes. There's the "HTML" syntax, which is basically tag soup, but at least there's an actual parsing spec these days. There's also apparently still an XML syntax for HTML 5? But the HTML syntax is the one that's common on the web.
> Missing from the spec is a formal model for tag inference and for deciding ambiguousness of content models. That void was quickly filled by Brüggemann-Kleins "One-unambiguous grammars" paper ca 1993, plus clarifications by James Clark
As far as I can tell, this sort complex and opaque documentation is one reason why promising standards get outcompeted by clear, simple 30-page specs. I've seen so many interesting technologies wind up as obscure footnotes in history because few companies could implement them at an acceptable cost. (Of course, today, a good open source implementation can play a much larger role.)
But to respond to your original remarks, I think this one of the major reasons that "each comp sci generation" ignores so much of the past. The past was frequently frustrating. It sometimes had glaring drawbacks. Often it was expensive and proprietary. And sometimes the ideas were too far ahead of their time.
I fondly remember at least a dozen technologies that contained brilliant ideas, but which never succeeded in becoming popular. Usually the reasons why they failed (or at least remained obscure) are obvious in retrospect.