Harmful Consequences of Postel's Maxim
tools.ietf.org
tools.ietf.org
On the first error, a browser should display an error bar in the middle of the page, then proceed on a best-effort basis, ignoring most styling. Bad pages would still be readable in emergencies, but they'd be annoying enough to get fixed.
(Until you write a web crawler, you don't realize how bad HTML is in the wild. There are low-level syntax errors. There are pages with more than one <html>, <head>, or <body> section, or where those sections are out of order. There are nesting errors which result in page trees a thousand deep. There are tags that don't belong in HTML. There's syntactically incorrect ad code from major vendors.)
Edit: I don't mean "draconian" in a pejorative sense, it is just typically the word used to describe this.
Graceful error handling is one of the things that set good applications apart from (really) bad ones.
When an SSL server sends just one bit wrong, that will lead to the connection simply aborting. Have you ever seen that being a problem for users? No, because servers simply implement the frickin spec when deviating from the spec means that your connection gets dropped.
You seem to have forgotten what HTML was created for. It's a document format.
If a browser did strictly enforce (what, HTML 4.01?) syntax, there would be no problem for users at all, because that browser would not have any users.
Now imagine most of those users don't know what a "compiler" is and are using the one that came with their OS.
Now imagine you are a competing compiler vendor who worked very hard to get some of those same users to install your compiler specifically.
And by the way, users expect the same compiler to compile all the projects and would be shocked and annoyed if they had to use different compilers for different code.
But HTML isn't a programming language, it's a markup language, and you absolutely can render a partially-correct page.
This might be more obvious if you consider a different markup language instead. Let's say you're using Markdown, and you start an italic block with an asterisk and forget to terminate it. Should the whole comment fail to render, or should it just fall back to showing a partially-correct comment?
But you absolutely can compile a program with syntax errors. There is nothing that prevents compiler writers to be equally "forgiving" and just inventing random semantics for stuff that does not have any meaning in the language, just as browsers do with HTML. And it would be just as idiotic.
Tell me, what do you think is the advantage of inventing some random meaning for what is declared to be an HTML document but actually isn't (bonus: that would not equally apply to code in a programming language)?
Secondly, it's pretty obvious what the benefit is of doing the best job you can rendering a document that contains some error in part of it. The existence of an error in one section of a document shouldn't have any impact on the rest of the document, and so there's no reason why the error should prevent rendering the rest of the document.
It seems like both views are quite defensible, depending on your theory of what the markup was meant for.
The view formulated in a non-misleading way is simply "if the standard does not specify semantics, I can as a matter of fact just use my own proprietary interpretation instead". Which is trivially true, but also completely useless if your goal is interoperability, while, if you don't care about interoperability, why care about the standard in the first place?
You are calling "indefensible" the way the entire industry actually works. Maybe you want to step back and look at why that happens.
2. What is even your point? That the industry has an outstanding track record of interoperability and security? What exactly is it that makes the view defensible in your eyes?
2. Standards are a tool to improve interoperability.
When standards are disconnected from reality and cannot be implemented by any actual players in the market, interoperability is harmed. It is not even neutral.
The kind of attitude you are showing here is what got us the idiocy of XHTML2. That is, many of the people who cared most about standards and semantics on the Web went and threw their collective efforts into a black hole for years because they were too blindly idealistic to engage with or even acknowledge the actual industry, that, in their arrogance, they thought they were standardizing. Meanwhile everyone else happily ignored them and all their efforts had no real-world impact on anything, except that by taking attention away from other, more realistic, standardization efforts, they probably set real interoperability back several years.
Reverse engineering and reimplementing the semantics of another vendor is obviously very much the opposite of "inventing random meaning", but rather "implementing a de-facto standard", and as such obviously helps interoperability.
But that in no way justifies the first vendor inventing that de-facto standard in the first place, whithout wich there would be no need for anyone to reverse-engineer and reimplement it.
If all deviations by browser implementors from the formal standard were reimplementations of deviations by other vendors for the sake of better interoperability ... there would not be any deviations at all.
Well, yes, you have. The definition of an error in this context is that there is no meaning for the given syntax specified in the standard. If you interpret the document anyway, you are assigning it some meaning. As that meaning does not come from the standard, you have simply invented it.
> The existence of an error in one section of a document shouldn't have any impact on the rest of the document, and so there's no reason why the error should prevent rendering the rest of the document.
1. You have it all backwards. Whether something is a "section" is part of the semantics that are specified by the standard. When your document does not comply to the standard, the standard does not assign any semantics to it, and as such it also does not specify what a section is in your non-complying document. The moment you talk about "sections" in your non-HTML document, say, you have already started inventing semantics.
Also, even if we assume that we somehow can magically identify "sections" anyway: How do you justify the assumption that sections do not have any impact on one another? For one, if you take HTML, in particular when CSS is involved, you absolutely easily can have very global effects using rather local constructs, can't you? But also: What if the human-language content has cross-references between sections? How do you ensure that the meaning of the human-readable content is conveyed correctly when somehow processing a "broken section"? What if the result is that you leave out an important warning regarding the following instructions in an attempt to help the user?
Really, how is that any different from a compiler replacing functions with errors with a No-Op?
2. What is the actual advantage? You say it is obvious, but it's not to me, which is why I asked. What is the advantage of just displaying something, even if the standard does not assign it any meaning, over aborting with an error?
The developer of the HTML markup must not rely on the browser doing "its best" with bad HTML.
During development, it's helpful if the implementation finds as many errors as possible in one pass and flags them all. Whether or not anything is rendered is immaterial; the HTML writer should fix everything.
The end user shouldn't see any errors, because they all blew up in the HTML writer's face and were fixed.
It is a fact that rendering happens in parallel with downloading a web page. So yes, renderers should be able to render a prefix of a web page, up to the first error, because the situation is entirely possible that only that prefix has been downloaded and the error is not yet seen. Renderers should not delay rendering until the entire page is seen; something has to be rendered when the prefix is available, but the erroneous text which immediately follows has not yet been seen.
Beyond that, no. Stop at the first error and splash a big red diagnostic into the output, so the user can see that the HTML producer screwed up.
Then when the HTML is tried with other browsers, it looks like crap.
So then the advantage is that it looks like other browsers trying to compete with yours are crap.
The competing web browsers developers are unhappy and grumble a lot, but in the end have no choice by to try to emulate the random meaning that you have assigned to the HTML. Of course, you haven't publicly documented it, so they have to spend time reverse-engineering it, which means they will be perpetually lagging behind.
Well, pretty much status quo in the world of ISO C.
[1] http://searchengineland.com/google-bans-iframes-for-adsense-...
HTML5 began as a splinter effort to write down what browsers actually do and mix in enough wishlists of what browsers should do, so it's only natural that they decided to codify many rules of handling broken HTML. While I too question whether it was the smart approach, it was definitely an approach that made sense to the stakeholders and authors of HTML5, who were largely actual vendors of browsers and parsers.
Even today, that browser vendors have pushed out all sorts of measures like near-mandatory TLS, stricter parsing of HTML is a goal no one really chases.
As an appropriate throwback, see Jeff Atwood on this topic in 2007 [1], complete with screenshots of different styles of browser error handling.
[1] https://blog.codinghorror.com/javascript-and-html-forgivenes...
Funnily, this wouldn't work anymore. The efforts to fix HTML have made it even more complicated. Now browsers insert a lot of fictious tags if you nest things improperly, and you have to simulate that.
I presume you mean html5lib and not the new html5-parser; html5lib isn't slow because of error recovery: html5lib is slow because it's a parser written in Python. The parsing performance is largely taken up, last I benchmarked the VM, by allocation overhead (heck, reading s[0] of a string s causes an allocation) and super-simple VM instructions (primarily dispatch overhead).
What I'd like to see is how much faster an HTML parser would be if it rejected any document with a parse error at the first parse error; I doubt it'd be much (you'd still need the branches to detect the parse error, though the code might be smaller and you'd get better cache locality).
I mean, you even could add deliberate delays, simply to incentivise site operators, because they tend to care about even relatively small effects on conversion rates and the like, while for most end users some sluggishness on some bad websites probably would not be enough to cause switching to a different browser?!
Your idea wouldn't work because it would never create any noticeable impact on page loading speed even in the best case scenario. That's assuming that browsers want to maintain and ship two totally different parsers, which (cf. XHTML) they don't.
This is likely true for most sites on the internet. It is worth noting that on an overly large HTML file of ~1.4 MB, stuffing that into an HTML element object by assigning innerHTML with that content takes approximately 75 ms on my i7-6700HQ.[1] A much larger file of 10 MB took 656 ms to parse, which seems to indicate it scales linearly and processes ~16KB/ms (16MB/s). For the most heavy real sites I could find, it generally took no longer than 25 ms.
It's worth noting that 16MB/s is actually only slightly faster than 100Mb/s, so some people may actually take longer parsing than receiving the HTML. That said, the parsing/processing tested here may be doing more than what we strictly care about for this diuscussion, so I'm not including it as a rebuttal, but as a point of reference that thought was interesting.
1: var el = document.createElement('html'); console.log((new Date()).getTime()); el.innerHTML = htmlstr; console.log((new Date()).getTime());
And, yes, it is indeed a sorry state.
Instead, relaxed or permissive parsing of files/protocols/API even though they have well-documented specifications is an unavoidable emergent phenomenon. The various examples across domains is fascinating:
- HTML (missing tags, unpaired tags, etc). Developers may prefer to strictly parse HTML and any non-compliance results in a "page error". But the websurfers want that page data so developers end up guessing the HTML authors intentions and renders the page.
- PDF (malformed pdf files that Adobe can read). Lots of utilities out there create bad pdf files and Adobe Acrobat has lots of workarounds to read them. Even though Adobe Inc controls the pdf specification, even they succumb to the pressure of adding workarounds to their parser to render broken pdf files. To add to the insanity, all 3rd-party industrial-strength pdf parsers end up copying Adobe Acrobat's behavior to parse broken pdf files!
- Win32 API. Programmers out in the wild will notice an undocumented behavior of an API that's not in the contract and start to depend on it. When a new version of Windows is released that removes the undocumented behavior and breaks the 3rd-party app, the customer blames the "Microsoft Windows upgrade" and not Quicken or videogame company. Raymond Chen has written several articles on Microsoft adding "shims" to Windows codebase to help "guarantee compatibility" for misbehaving apps incorrectly using the Win32 API. If the 3rd-party app is important enough to consumers that it prevents them from upgrading Windows (in other words -- pay Microsoft money), it means Microsoft will bend over backwards to accommodate the badly-written code from the vendor.
The external forces are too great to enforce perfect discipline of strict parsing across all parties. Even the big proprietary companies like Microsoft and Adobe can't enforce their own standards-compliant parsing. An open standard from IETF would have no chance at all.
We'd like to think of a "computing standard" as some inviolable contract but history has shown it is actually an organic (and unspoken) "social" agreement. This is unavoidable unless we all agree to have all software "approved" by a central authority before anyone can download or use it.
Note that similar thing was happening with SSL/TLS connections for a long time: developers may prefer to strictly verify X.509 certificates, but websurfers want that page data. The result of developers bowing to accepting invalid certificates was hurting everybody.
Nowadays the situation is much better, because browser developers made visiting sites with self-signed, expired, or incorrectly named certificates significantly harder.
I categorize that scenario in a different bucket because novice users don't understand the security implications of broken TLS. So, the users don't want that data but they don't know it. The developers in this case are looking out for the user. (Same security situation as developers helping the user by restricting cross-site scripting or address bar hijacking.)
To me, that's not the same as the forgiving parsers of HTML and PDF. Whether the HTML is missing </p> closing tags like this:
<p>This is paragraph 1.
<p>This is paragraph 2.
Or is 100% compliant like this: <p>This is paragraph 1.</p>
<p>This is paragraph 2.</p>
... the websurfer just wants the page rendered in both cases.Oh, of course it's not the same. The point is, in both cases users want strictness, but don't know that they want. Users want web pages that are easy for their browsers to parse; "want" to different degree, yes, even an indirect "want", but still.
If browsers fail to display web page with invalid markup, developers are forced to fix their sh&t instead of adding another exception to an already big pile of "steaming let's not" in the standard. And we had a precedent of going down this way with X.509 certificates. I think it's a pity that the way is too hard for something not as critical as SSL/TLS.
Compilers solve this by having an intermediate levels of strictness: warnings, which a developer who's providing input to the compiler is expected to fix, but which an end-user doesn't have to. This captures the benefits of both strict and lax input checking.
It helps to make similar distinctions in other contexts, whether that's having an explicit notion of "warnings", or just putting messages in log-file output. It also helps if a format has a well-known extra-strict checker around, which developers can test their programs' output on.
No, you absolutely don't. The only reason why you would possibly need that in the first place is because software tends to be far too forgiving. Where software enforces procotol/format compliance, a normal end user essentially does not ever get to see any deviating input because generating software is written to the standard right from the start, and the few cases where that fails, the actual bug gets fixed.
Also, you actually can't. When the standard does not specify the meaning of some input, then there is no meaning. Whatever meaning you invent for it is just that: Your invention. The next software likely will interpret things differently. You do not actually have a standardized format when multiple implementations speak the same syntax, they also need to have the same semantics. Just doing something in response to broken input does not produce interoperability, but just undefined behaviour.
https://www.reddit.com/r/programming/comments/6u1jq2/the_har...
Some refuse empty fields and demand they not be included as tags. Some demand a specific list of tags, and insist all be included even if empty. Some clearly have non-recursive parsing, only honor certain depths of tree. Some have completely hard-coded parsers, and only accept certain tags in certain orders. The list goes on.
Is this just Sturgeon's Law, where things are broken in both directions? Or is there some predictable reason that some tools become overly permissive, and others overly restrictive?
In every language I work in, I find myself writing a lenient JSON parser. So far, Python, C++, and Nim (for JS, there is already JSON5).
I am also working on a whitespace and comment preserving parser that round-trips, so you can change a value in a configuration file without loosing comments. (Currently, I always produce standard JSON so compatibility is not an issue.)
I have been thinking about publishing these parsers, but after seeing how much abuse and nasty comments the JSON5 guy got, I have been holding off...
I am 80% done writing such a parser for a over-engineered legacy config file format. It isn't that hard really, you just need a parser framework/library that can output a complete parse tree (a tree where a preorder traversal covers every byte in the input contiguously, exactly once).
The remaining 20% of the work is keeping track of the line and column number (almost there), cleaning up the code, and being able to reserialize the AST to a binary format (for reloading in another process where I don't want to have to rerun the original PEG)
If we did not have Postel's Maxim, I think the Internet would have just remained the domain of geeks and not have become what it is today.
> If we did not have Postel's Maxim, I think the Internet would have just remained the domain of geeks and not have become what it is today.
That would have been a huge benefit as well.
However, the IETF is moving to markup as the authoritative source, so the pseudo plain text version might be losing relevance.
But they have included an 'html' option in the top menu which, if you click, takes you to the same document with standard webpage format.
I guess they could make that the default and then provide the current default as a 'printable html' option.
Jon's principle could perhaps be more accurately stated as 'In general, only a subset of a protocol is actually used in real life. So, you should be conservative and only generate that subset. However, you should also be liberal and accept everything that the protocol permits, even if it appears that nobody will ever use it.'"
Edit: actually, some of the replies in the thread you linked disagree with this remembrance, and state that it really did mean what we take it to mean today.
The context for the various iterations of Postel's Law generally suggest that they're referring to the possibility that you might be seeing the result of mismatching versions of a specification. The later iterations also give reference to an explicit example of what they mean: don't assume that enumerations in the specification are closed (i.e., assume that future revisions may add additional enumerations). There is absolutely no evidence that he is advocating trying to parse slop at all.