(Consider: Apparently, MS-Word docs or PDF prove longer lived than basic HTML documents! Who would have thought of this?)
(Consider: Apparently, MS-Word docs or PDF prove longer lived than basic HTML documents! Who would have thought of this?)
- https://whatwg.org/working-mode#removals
- https://whatwg.org/faq#removing-bad-ideas
In short, I think it's important to distinguish between conformance and removal from browsers. Removal from browsers is a big deal and, as per those links, is only done when it's not going to break the web, or when the benefits are very high (e.g. security issues). Removal from being conformant just reflects the evolution of best practices. See also https://github.com/whatwg/html/blob/master/FAQ.md#how-are-de...
Removing support for presentational markup does not mean a loss of information. Browsers will still render tags they don't recognize, and re-applying the styling of those tags is often trivial. (I mention <blink> because it's one of the more difficult, but not terribly so)
As the web evolved, our needs changed. Do we still need the <font> tag?
It's true that you can just use some CSS to make up for the lost HTML feature, but than again you could also rewrite the HTML part.
Forgive me if I'm wrong, but I'm fairly sure that what the OP is trying to say, is that there are plenty of great websites out there, which were developed a long time ago, and for which there is no maintainer to do any work on it. Thus having HTML elements like this dropped, would make the content in a way lost.
Thinking about it some more, users can probably add plugins to add this css automatically, or some browsers might even keep those features in, but still, there will be users that don't know this, I think, resulting in a bad experience.
The content will not be lost. The tags will result in valid elements but the rendering may vary. This has always been a thing to be expected, since legacy elements (pre HTML5) never had uniform rendering and contained quirks.
Should the current/new standard have support for ambiguously rendered quirky elements? Is it even a standard then?
After HTML5 the end result will definitely be the same on most (if not all) layout engines. The standardization as a process requires non-conforming legacy to be dropped.
I was thinking about something like Stylish or the user stylesheet I've been hearing about in Firefox (for their UI IIRC, but still). Inject some global css on older/missing doctypes, and it's probably less than 200 total declarations to handle every older tag. I'd imagine <font> to be the hardest and/or longest, followed by <blink> and <marquee>.
Would be a small extension.
My other point I think I expressed clearly enough, that the loss of presentational markup is not a loss of content in most cases. If the title is in Times New Roman instead of Arial, most of the time it'll just look worse. Unless the content is meta, the presentation is to make things more pleasant to read.
My primary point is that presentational markup is not required to get value from all but the most meta of old pages.
<span style="color: #000; font-face: Whatever; font-size: 10pt">Blah</span>
vs. <font face="Whatever" size="2" color="#000000">Blah</font>You could use "font-size: small" to get the size="2" behavior.
(Also, "font-family: Whatever", not "font-face: Whatever".)
> <span style="font-size: 10pt">Blah</span>
vs.
> <font size="2">Blah</font>
Which is easier to remember? Which is more obvious at a glance?
Plus, the former is an absolute value where the latter is not, as far as I've been able to tell. I know that first span will always be 10pt font. I have no idea what "2" even means in this context.
HTML was never intended to be archival. Archival assumes a long term relationship between format and user-agent, but those two things evolve independently.
> Who is going to update these documents in order to make them conforming to future browsers?
You don't update legacy documents stored in an archive. You find a conforming user-agent (appropriately old browser version) to consume them in their intended state.
> Is it worth it?
Yes, HTML is a versioned format. Improvements to the format are welcomed and necessary.
From what I understand, the WHATWG's policy regarding archival is "well yes the format is constantly changing but we'll try REALLY hard to not make too many breaking changes."
That said, validity changes don’t matter to the browser’s ability to render old pages. Changes to remove support for an element entirely are very rare.
Additionally, WHATWG lost some credibility when they attempted to redefine the DOM and arbitrarily delete some node types. Granted, most of those types are legacy types not in use by anybody in long time, except for the attribute node type. Browser vendors simply ignored this foolishness.
Here is a very simplified description of this problem years after the fact: https://github.com/whatwg/dom/issues/102
It is important to understand the DOM wasn't created for HTML. The DOM, starting with DOM level 2, was created in parallel with XML Schema. This is evident when reading some of the W3C mailing lists and comparing release dates of W3C publications.
Attribute nodes can be independently walked when walking the DOM. By removing attributes as a node type you break this functionality. You can use this little utility I wrote as a proof: https://github.com/prettydiff/getNodesByType/blob/master/get...
Browser vendors are extremely shy about adopting new technology that makes for breaking changes. They will do so, but you need to have an incredibly strong argument. WHATWG's changes to the DOM had no beneficial argument, except perhaps developer convenience for those developers who cannot figure out DOM walking.
The DOM is a pretty solid technology with regard to extensibility, predictability, and sturdiness. If you maintain a large major browser and somebody came to you with breaking changes and a bunch of weak bullshit for justifications what would you do? Also, imagine if you will, that if you ever challenge the people bringing you this pile of shit they will troll the hell out of you in a very visible and immature way.
The response from the browser vendors was to simply say nothing and ignore them like they were never there. I got into an argument about this with the WHATWG on a github issue once, and wish I hadn't. Ignorance is like a black hole that sucks everything in and it never stops to allow rational signals to escape undamaged.
Regardless of issues like this, browsers track WHATWG DOM near exclusively. You can see devs from all of the major browser engines commenting in the issue you linked.
I know from my own conversations with the WHATWG this wasn't something that long time WHATWG members would admit to (or even understand). It was the childishness, perhaps more than anything else, that nobody took them seriously.
> Regardless of issues like this, browsers track WHATWG DOM near exclusively.
I am going to disagree with you there. Perhaps they do now, extremely recently, but historically this is absolutely false.
> You can see devs from all of the major browser engines commenting in the issue you linked.
Yes, everybody participates in the WHATWG. This isn't new. Participation is different than adopting those recommendations back into your software.
Here is what browsers actually implement: https://www.w3.org/DOM/DOMTR and https://www.w3.org/TR/dom41/
It is important to keep in mind that the WHATWG doesn't do a lot of XML work, but the DOM is markup language agnostic. The DOM isn't something created or maintained in an HTML rich vacuum.
The person who ultimately fixed this problem in the DOM Living Standard is Anne Van Kesteren, who was not even remotely new to WHATWG at the time. The person who filed this issue (Philip) is also a WHATWG old timer.
Opposed to this, HTML was not intended as presentation layer for fancy web-apps. (There had been better around for this in the Hypertext-world, even then.)
This is about the exact opposite of archival: backwards compatibility. We don’t want to split the web into old web and new web. Having to switch browsers for decade old pages as we encounter them, raising the barrier to entry for that lore of old, effectively sepulchring it from the public.
99% of the web’s users are not going to understand when to switch browsers, how, nor why.
It happens anyways regardless of what people want. The 90s era web doesn't work properly in modern browsers and 90s era browsers don't work with the modern web.
> 99% of the web’s users are not going to understand when to switch browsers, how, nor why.
This also happens naturally. Chrome is the most popular browser and it doesn't come with most operating systems. That is something users must switch to.
The comment mentions long-lived MS-Word doc. Which version of those, specifically, has been around the longest so far?
As for MS-Word and HTML: When I finished my thesis in the mid-1990s, I saved it both in the MS-Word version, I used to write it (MS Word 5 for Mac), and HTML (expecting future compatibility). I can still open the Word version, but I may be soon unable to conjure a formatted display of the HTML-version. And I can still display a PDF 1.x...
(Current frame is "self" or "window", parent frame or frameset is "parent", and the top most entry point into the hierarchy "top". Moreover, "self", transcending the window context, is also the only reliable reference to the global object, thus also providing a valid reference to the context of a worker. Specifically, it was for framesets that the notion of hierarchy was introduced, which eventually resulted in the concept of the DOM. Some inconsistencies to this concept of strict parent-child relations were actually introduced by early implementations of the iframe-element, which is, BTW, still a valid HTML element.)
That said, there was a small inconsistency with an early subversion of Netscape 3, regarding, whether the frame source would be relative to any current location of the frame or rather relative to the frameset. (But this was an issue for a rather short period of time, two months or so.) A major difference in styling was the implementation of frame borders, if they would be entirely invisible by just specifying `border="0"` (Netscape and others) or, if they required the two attributes `frameborder="0"` and `framespacing="0"` (MS IE). In practice, next to all sites specified both schemes. And jet another, but minor implementation specific detail was the sizing of framesets: While Netscape Navigator supported, like all other browsers, a size specified in pixels, this was internally translated to percents of the total width. Therefor, depending on rounding to integers, the presentation in the Netscape browser could be off by a pixel or two.
(The latter was, in deed, not unusual behavior at the time, just like MS Word and RTF used to translate any measurements internally to "tips" or twentieths of a point.)
I am super glad for all the hard work put into all this.
The report is saved across six .doc files, due to the size limitation of the 3.5" floppy disks we were using back then.
Anyway what's going on anyway with Google+Microsoft+Apple+W3C, why is there such a big push to HTTPS and HTTP/2, and declaring old HTTP/0.9 and HTTP/1 and HTTPS/1 and HTML5.0 as legacy!? And why is Mail still sent in plain text completely insecure, and no adoption hype to support SMIME/etc? It is beyond fishy. Or is it just pure greed, no one cares about non-walled-garden-open-web (aka everything has to live in LinkedIn/FB/AppStore/PWA) and there is no money in mail?
The actual HTML specification which browsers follow is maintained by WHATWG at:
...is found here: https://chromium.googlesource.com/
The reality is that the WHATWG (a) only writes descriptive standards, describing what already exists, usually with pseudocode and prosa instead of ABNF or EBNF (see the URL standard replacement), and (b) only describes something once it’s actually been implemented on larger scale.
On the topic of what standards are supposed to do – prescriptively shape and replace what exists – the WHATWG isn’t useful. WHATWG "standards" are the equivalent of Microsoft Office Open XML, a standards body just taking an existing implementation, defining whatever it does as standard, and doing it so incomplete that the result is useless.
Yes, WHATWG and W3C are doing the best they can do in the current climate (where Google can roll out QUIC and SPDY before even any standard is defined across websites accounting for 6% of global traffic, 65%+ of web browsers, and 85%+ of mobile phones), but this is just misleading. It helps no one to pretend to do standardization work when you don’t actually have any power to decide anything – neither WHATWG nor W3C can actually force, or even ask, Google to change SPDY or QUIC. They’re papertigers.
In Chrome we ensure that all features we ship to the web go through a public standards process. This allows them to be developed by a collaborative community, including other browser vendors and web developers who would use them. It ensures that if we happen to ship a feature sooner than other vendors, there's a specification and a shared test suite (https://github.com/w3c/web-platform-tests) that allow others to quickly follow. Note that a specification is better than requiring them to read the Chromium source, because specifications are at a higher level that doesn't depend on individual browser architecture details.
In the WHATWG we don't only write descriptive standards. But we do ensure that whatever standards we write, are ones browsers are willing to implement. And we ensure that standards accurately describe how browsers operate, even for legacy features, because that is all part of the mission of allowing browsers to compete on an even playing field and build themselves from scratch without having to go through the kind of costly reverse-engineering that Firefox 1.0 did to catch up to IE6. In practice we've found that algorithmic specs are better for this than BNFs, as it's harder to specify error-handling behavior for BNFs while still staying compatible with the web (i.e. while still producing a standard browsers are willing to ship).
And yes, we're not interested in just creating a standard out of thin air, with no vendor collaboration, calling it "standard", and then hoping some magical power would force browsers to implement. It is indeed much more collaborative than that.
But the fact that we require standards to be developed in tandem with implementations doesn't mean that implementations (such as Chrome) just go ahead and do whatever they want, and we at the WHATWG transcribe it into the spec at some lower level of detail. Instead, the public, collaborative standards process helps to extract out all testable and observable aspects of the feature into a codebase-agnostic description others can use, and provides a forum for them to comment on ideas before any final shipping decisions are made. And, per our working mode (https://whatwg.org/working-mode#changes), changes and additions do require multi-implementer support before they're ready to graduate to a WHATWG Living Standard; proposals not yet at that point are said to be in incubation, and are often developed elsewhere (see https://whatwg.org/working-mode#new-proposals) such as the W3C's WICG.
Ehm, basically every major feature Chrome has shipped has been shipped before the standard was even discussed. SPDY shipped long before HTTP/2 was even finalized, and QUIC is doing the same. NaCl shipped in the same way, without any standardization, and to this date, earth.google.com depends on it.
In general, your problem is that you only consider browser developers. In the past, the WHATWG has decided to redefine the URL standard, then shame cURL for not following the standard, without ever involving anyone from the curl project in the discussion. The URL discussion affects everything from Android’s IPC system to curl, from industrial machinery to the web. The WHATWG explicitly declared that the URL spec is designed to completely, and exhaustively, obsolete and deprecate any existing URL or URI spec.
Yet, the only people ever contacted about this, and who were given the ability to take part in the discussion, were representatives from the three large browser vendors.
The URL Standard was designed in the open with input from many different constituencies. The cURL author has chosen not to participate, for reasons of his own, but e.g. Node.js, PHP, Google's GURL (used by Android IPC, I believe), and others are quite involved.
The only participants were all either browsers, affiliated with browsers, or a handful of web serving projects.
Other projects that rely on URLs include everything from KDE to Gnome, Microsoft’s OS to the systems used in your car.
Changing a URL standard and only involving web vendors is basically like changing the A4 paper standard and only talking to the Microsoft Office team, the Google Docs team, and HP’s printer team – while entirely ignoring paper manufacturers, envelope manufacturers, the mail companies around the world that will have to ship the envelopes, fax manufacturers that have to build faxes able to fax the new format, newspapers and magazines that have to replace their paper, newspaper shelf manufacturers that build newspaper shelves for newspaper stores, etc.
Most of the time, it’s easy to only think of the web as browsers and servers, but some of the specs the WHATWG touches go through entire industries, sometimes there are millions of companies that have to be notified months or years beforehand to replace their software, update it, potentially even do a recall, and standardize. Not everything moves as fast as the web.
And this entirely disregards the people that are trying to parse the web with HTML parsing, which everyone loves to ignore. And so many other groups of people and companies.