The havoc the XML purists wrought
scripting.com
scripting.com
He talks as if JSON is somehow "hard" for developers to grok. It's about as simple as you can get. I don't think it's noble to support a subset of XML, where regular JSON would fit better, merely to placate some developers used to doing things the hard way. If anything, API developers should be leading us by the nose to do things the best way.
Even if you "ignore" much of XML, as Winer suggests, you can still stumble into trouble. I've written (and used) a few RSS parsers in my time and dealing with broken XML with a regular XML parser is a gigantic pain in the ass. Many developers don't sanitize their input properly or dump source HTML or XHTML verbatim into their <description> elements. Try throwing that through Expat without thinking about it.. You can screw up JSON too, of course, but it's almost entirely down to quote marks alone.. not tagging, erroneous character encodings, non-existence namespaces, and more.
In contrast, Amazon supports XML in their web services, quietly and competently. I can't imagine them saying one day "It's too much work for us to keep supporting XML so you all have to rewrite your code now if you want to keep paying us for the web services you use."
Amazon's Web Services are good/unique enough for people to put up with the bullshit of XML in order to use it.. though their command line and Web interface tools are good enough that I suspect most users never need to get down to dealing with XML anyway.
A better comparison than RSS v. XMPP is RSS v. Atom. And in this comparison, RSS is by far the loser. It's incredibly difficult to write an RSS parser which can handle even a small fraction of the RSS published today, mostly because the standard is absolute garbage. In contrast, an Atom parser can be knocked out in an afternoon.
I'd be happy with a "reduced set" of XML, which excludes stuff like DTDs, named entities, and references. I've never seen these features put to any significant use in real life. But most of XML is quite sane (if a bit verbose).
What I don't get, with XML in the real world, is why people prefer this...
<A><B>1</B><C>2</C></A>
instead of my preference...
<A B="1" C="2" />
With the first form, you need to write out each tag name twice, and there's ambiguity when whitespace appears between tags. Yet that's what most XML out there looks like. Maybe I'm just odd,
EDIT: ok, the first link is to an IBM Research article which I recall having been useful: http://www.ibm.com/developerworks/xml/library/x-eleatt.html
The difference is much the same as that between native datatypes and objects. Some languages explicitly expose everything as objects, while others contend that, for example, boolean-objects are overkill and downright problematic.
Generally, I use attributes for strongly typed data, and elements for untyped or (semi)structured data. Schemas are vital for typing.
<a>
<b>hi</b>
</a>
is not the same as <a><b>hi</b></a>
But Erik Naggum makes the argument against XML so much more fun <pre><b>Hello!</b></pre>
is rendered differently from <pre>
<b>Hello!</b>
</pre>I, however, really enjoyed this line from the link. "I once believed that it would be very beneficial for our long-term information needs to adorn the text with as much meta-information as possible. I still believe that the world would be far better off if it had evolved standardized syntactic notations for time, location, proper names, language, etc, and that even prose text would be written in such a way that precision in these matters would not be sacrificed, but most people are so obsessively concerned with their immediate personal needs that anything that could be beneficial on a much larger scale have no chance of surviving."
This wonderfully, succinctly explains why efforts like the semantic web are doomed to failure.
an <strong>emphasized</strong> word
from an<strong>emphasized</strong>wordNot entirely relevant to your point, but worth bringing up. (XML doesn't ever mangle binaries, because it simply forbids them.)
Saying there are no s-expression parsers for non-Lisp is ridiculous.
This one's http://sexpr.sourceforge.net/ for C, it's been in sourceforge since 2002.
perhaps you prefer one in :
Javascript - http://planet.plt-scheme.org/package-source/dherman/javascri...
Perl - http://search.cpan.org/~nelhage/Data-SExpression-0.34/lib/Da...
Limbo - http://man.cat-v.org/inferno/6/sexprs
I shan't go on
But you'll note the disparity in support. A library for parsing XML with C may have existed since 2002, but it's existed for XML since 97. In Java, XML support has been baked into the JVM for a while, but S-exp parsers are hard to find. That library you list for Javascript is not a s-exp parser but rather a Scheme library for generating Javascript from a Scheme-like language. Googling for javascript and s-exps returns a lot of results like that but not many s-expression parsers. There certainly aren't any widely used and recognized libraries like there are for XML and JSON fro most languages aside from Lisp.
A lot of XML-hate comes from its overuse—it comes with overhead that is too expensive for simple cases, but is pretty useful when you go beyond simple. Right tool for the right job and all that...
JSON:
{"a":{"b":"hi"}}
YAML:
---
a:
b: hiIf a node doesn't mean anything to your application, ignore it. If you're an XHTML parser, the whitespace is significant, so you need to handle it. If you're a FooML parser, the whitespace is insignificant, so just ignore it.
If you hate XML, this is not a particularly good justification.