HTML Validation: Does It Matter?
codinghorror.com
codinghorror.com
Little did they know that Microsoft would come along and tear down that world. Today, we've seen it bounce back a little, but really there are three rendering engines - Trident, Gecko, and WebKit. And, compared to what we had back in the Netscape days, they're all good engines. Sure, we might consider Trident crap by comparison to Gecko and WebKit today, but it's clearly a cut above what we were seeing in the Netscape days.
I think the reason that people today don't care about validation is that perfectly valid code can render differently in the three different rendering engines. If valid HTML doesn't give you the benefit of proper rendering, then what value does it give you? If a rendering engine is "wrong" in how it renders your page, you can bet that's your problem and not the rendering engine's problem - at least it will be according to your visitors. And three rendering engines aren't too hard to test - especially given the similarity in a lot of the rendering between WebKit and Gecko.
As for XHTML and HTML, Apple's "Surfin' Safari" blog has a great post on it: http://webkit.org/blog/68/understanding-html-xml-and-xhtml/. I guess XHTML never really fulfilled its purpose as having machine readable data in HTML - possibly eclipsed by things like RSS.
This is exactly true and the reason is that web standards aren't real standards by the textbook definition of the word, mainly because they lack what's called a "reference implementation".
I forget where I read it first but the overall gist is that in other fields, say industrial design, a standard typically isn't ratified as finished until the governing body creates an example for everyone else to refer to, the reference implementation, rather than just relying on the written descriptions of how the thing should work. For example, Firewire has a reference implementation that's certified by the IEEE and that hardware makers would refer to when building out the first versions of it. You could grab the example version, plug it in the way it's supposed to, test out the data that's flowing through it, measure what speeds its capable of, etc.
Since web standards don't have that reference implementation by the W3C, browser makers are left to their own devices to parse the insanely complicated technical language on their own. And it's only getting worse - the html5 spec is something like 10x the amount of writing as the last version of html. This is one of the biggest reasons why rendering engines have effectively stagnated in the last few years - things are "good enough" that they can wait around until the spec makes sense and there's working examples in the real world, and/or they are spending a gazillion hours trying to figure out what the new specs actually mean.
This is the problem with functional specs that 37signals rails against when they say that "Functional specs only lead to an illusion of agreement":
http://gettingreal.37signals.com/ch11_Theres_Nothing_Functio...
This is a super sad state of affairs, and not just in the theoretical sense of how fast the pace of development could be for browser makers today if things were more streamlined, it's also a huge hindrance for developers in the trenches like me. I haven't thought about this for years but I used to relish reading the CSS spec and exploring new ways of pushing my work forward. Trying to read the W3C docs now is like speed reading pig latin - you know there's some cool stuff in there but it hurts your head just to even try hunting for it.
I'm not going to go into details of why he's wrong, because a) this has been discussed to death by people who actually matter in this space, and b) there's a ton of comments on his post explaining what's wrong with his view.
There are ways to argue that validation is unnecessary; however, they do not include: 1) bringing up examples such as "target" attribute, and 2) saying "who cares if it works anyway."
His argument about user-generated content is also spurious. If you dont parse and filter user-generated content (with a real HTML-parser, not just regexes), then you are vulnerable to script-injection attack.
Jeff is a respected blogger.
"Google actually ranks it's indexed pages. The more valid the (X)HTML of your pages, the higher it'll appear in a search."
"If you do write XHTML, you'd better get it right. I heard about CodeProject practically dropping off the map because of an XHTML error that caused Google to stop ranking them."
Is this indeed documented or well confirmed behavior of Google? That might be a pretty important reason to validate.
Anymore I actually like validating, I use it to check to make sure nothing is blatantly wrong. There are a few rules that do not make much sense, but overall I find it very helpful. But to make use of it this way don't go and build the whole site, then try to fix all the problems and bad decisions made earlier, use it through out the design process.
I know this blog post comes to this same conclusion by the end, but it just feels so wimpy and pandering to everyone. Through the whole post it goes from validation is pointless, painful, useful, helpful... do whatever you want. Maybe the author should just take a stand and stop trying not to offend anyone.
Error handling is not defined in the standards so if you rely on it you're at the mercy of every idiosyncratic error handling decision of every idiosyncratic browser implementor.
It is difficult to write a whole website, validate it and fix the errors and at this point validating has little value. However by validating periodically you are getting more out of the tools available. I don't force myself to write 100% strict HTML, but minimizing the validation errors means that I will have less to read when it's time to.
With that in mind, though, it's easy to write valid markup. All the rest of the code in your app has to validate, why negelct the HTML? Making sure that your markup is valid will save you time when you have to debug a weird rendering issue (or JavaScript's view of the DOM).
Finally, if you even have the opportunity to pass invalid markup from your app to the browser, you are probably doing something wrong. In my apps, I load the web designer's templates (which may be invalid HTML) with libxml2 (it has an HTML loader), process them programatically, and then serialize the DOM for the browser. This way, the browser always gets something valid -- my program's type system enforces this.
Valid HTML means you have one less thing to worry about.
A couple of days ago I built a new webpage and as usual tried to validate my HTML 4 strict and CSS 3 . The HTML4 validated just great with CSS2.1, but threw a tantrum when I posted the CSS3 validation code and reran HTML4 Strict.
This is the offending line:
&profile=css3&usermedium=all&warning=1
which is OK on the CSS validation page, but will not pass inspection if ran AFTER posting the WHOLE link to my web page. Everything is the same except for that last line.
Run as CSS2.1 and all is well.
Run as CSS3 and the validator gives 6 errors on this part of the code alone. It seems to hate the " & " and the " = " and even the " W " in warning?? !!
Would like to validate HTML 4.01 Strict with CSS3.
Any help??
When I develop I use a firefox addon http://users.skynet.be/mgueury/mozilla/ that automatically tells me how many validation errors I have (it runs the validation locally).
So as much as possible my code validates, but sometimes it doesn't - but each time, there is a reason for it, and it's not just an accidental mistake.