What is XML good at? (by good I mean better than alternatives like JSON, YAML, HAML, etc)
The only thing I that might qualify is a long term/archival quality document format like ODF/OOXML. The inherently embeddedable nature of XML does seem like a nice fit but it gets very bloated very fast (deflated wrappers help though).
Did you know that web browsers have native support for XML? Try fetching an xml resource and pulling the responseXML value off the XHR - you have another DOM object right in your hands, that you can treat just like a regular Node.
The reason that XML is better than the alternatives has nothing to do with the syntax itself, but the tooling around it. Browser support, XQuery, XPath, XSD/RelaxNG, XSLT - please tell me where I can find the equivalents for any other markup language. You don't have to use them, either - but if you need them, they are there in pretty much every framework. XML had the first to market advantage, and was picked up in enterprise systems and made powerful and ubiquitous. If you need batteries included, XML is right there, the others are not. There is really nothing wrong with it, the syntax is clunky for some applications but great for documents, whereas e.g. JSON would be horrible in that scenario.
Not only that, browsers have support for XSLT 1.0, so you can format it and style it with CSS on the fly. Even mobile browsers support that, as far as I can tell.
2. XML data model is more sophisticated than JSON or YAML: it supports element ordering and mixed content and does it rather elegantly and succinctly. It also has namespaces (and these are very good namespaces, they're not hierarchical, they're just long names in a single flat namespace with convenient notation to shorten the long prefixes to reasonable size). As a result it's very easy to define a new language, extend a language, mix multiple XML languages, etc. JSON and YAML are hopeless here.
3. XML comes with tools to define the type of the document or a fragment, so you can read a document and automatically check that it has the right syntax (and/or convert the data, such as dates, into the native format). There are three ways to do this (DTD, Schema, Relax NG) in order of increasing power and expressiveness (not just syntactic sugar, but different kinds of languages). In particular, it natively supports things like inter-element references, which is very convenient for complex documents.
4. XML comes with XSLT, which is a general-purpose tree transformer (transducer) with declarative syntax. This is an immensely valuable tool. To put things into perspective: a compiler is a special-purpose tree transformer that transforms the source tree of a program into machine code (which is also a tree, technically: sections, data, functions, etc.). Are you sure you don't need a general-purpose declarative tree transformer and prefer to write ad-hoc ones? :)
5. The specification of XML 1.0 is shorter than, say, YAML :) OK, this is only one part of XML landscape, the whole is much bigger, of course; but still this part (basic XML and DTD) is noticeably shorter than YAML. (I myself also find YAML pretty cryptic.)
Supporting XQuery/XPath queries and typing with DTDs are the big ones for me, including all the surrounding tooling. You can get replicas of both of these in JSON now, but I don't think they're as mature.
(For the purposes of this post, I'm including HTML in the XML family.)
XML/HTML is good when:
1. You have two dimensions of markup you want to do. That is, you have a clear distinction between what is a new "tag" and what is an attribute on that tag. If you can't almost instantly decide whether some feature you want to add works as an attribute or a tag, you probably shouldn't be in XML.
2. Almost every tag one way or another contains some text, the third dimension that XML supports. A proliferation of tags that never contain any text is a bad sign. A handful may not be a problem, e.g. "hr" in HTML, but they should be the exception.
3. You have a really good use case for XML namespacing, the fourth dimension of information that XML supports, in which case there's almost no competition for a well-standardized format, as long as you're also using the previous three dimensions.
There's sort of this popular myth that XML is useless, which I think isn't because it's true or that XML is bad, I think it's because in general, most times you want to dump out a data structure #1 isn't true, let alone #2 or #3. In a lot of data sets, you've only got the two dimensions of "simple structure" and "text", not annotations on the structure itself. (Or, perhaps even more accurately, they end up implicit in the format itself, and the format is constant enough for that to be just fine.) A lot of stuff in the 1990s and 200xs used XML "because XML" even though it clearly failed #1. XML is really klunky when you don't want that second dimension because the XML APIs generally can't let you ignore it, or they wouldn't actually be XML APIs.
On the other hand, when you learn this distinction, you do come across the occasional JSON-based format that clearly really ought to be XML instead. You can embed anything you want into JSON, but when you're manually embedding a second structure dimension into your JSON document, it loses its advantages over XML fast. If you've ever seen any of the various attempts to fully embed HTML into JSON, without leaving any features behind, you can begin to see why XML or XML-esque standards like HTML aren't a bad idea. HTML is much easier to read for humans than HTML-in-JSON-with-no-compromises.
And if you've truly got the four-dimensional use case, XML is really quite nice. When you need all the features, suddenly the libraries, completely standardized serialization, and XPath support and such are all actually convenient and surprisingly easy to use, for what you're getting.
Some examples: HTML is a generally good idea. SVG is a middling idea; it passes #1 and #3 but fails #2. SOAP and XML-RPC is generally a bad idea; SOAP fails #1 and #2 but sort of uses #3 and XML-RPC fails all three. XMPP I actually think is pretty solid as an XML format (mere network verbosity problems can be solved with an alternate encoding, though admittedly that becomes non-standard), and in a lot of ways, the real problem with XMPP isn't so much the format itself as that people are not used to dealing with the four-dimensional data structures that result. People expecting IRC-esque flat text are not expecting such detail. Using the fourth dimension of namespaces for extensibility is neat, but few developers understand it, or want to.
I generally don't just comment "attaboy" but there you go.
Not quite as bad as XML.. I think the problem is more the verbose, overly-nested format that was chosen for MathML than XML itself though.
{
"mrow": {
"mi": [ "x", "=" ],
"mfrac": {
"mrow": {
"mi": [ "−", "b", "±" ],
"msqrt": {
"mrow": {
"msup": {
"mi": [ "b", "2" ]
},
"mi": [ "−", "4ac" ]
}
}
},
"mi": "2a"
}
}
}Basically the problem is this: Either you explicitly represent the grouping in a general scheme capable of it, then you get the disaster (from the point of view of human readable and manipulable) that is XML or JSON. Or you use a domain specific language like LaTeX or whatever, with the attended parsing issues,etc.
If you want people to edit it by hand, the latter option is much better - but it has it's pain points. You don't get to use a broad range of robust tools to manipulate them, for one thing.
JSON is great at many things, but polymorphic substructures are AFAIK only really possible with everything being an object defining the "type" that it is. And that looks significantly uglier than what you have above:
{
"type": "mrow",
"children": [
{
"type": "mi",
"identifier": "x"
},
{
"type": "mo",
"operator": "="
},
{
"type": "mfrac",
"rows": [
{
"type": "mrow",
"children": [
{
"type": "mo",
"operator": "-"
},
{
"type": "mi",
"identifier": "b"
},
{
"type": "mo",
"operator": "±"
},
{
"type": "sqrt",
"expression": {
"type": "mrow",
"children": [
{
"type": "mi",
"identifier": "b"
},
{
"type": "msup",
"expression": {
"type": "mi",
"identifier": 2
}
},
{
"type": "mo",
"operator": "-"
},
{
"type": "mi",
"identifier": "4ac"
}
]
}
}
]
},
{
"type": "mi",
"identifier": "2a"
}
]
}
]
} { "mrow": [
{ "mi": "x" },
{ "mo": "=" },
{ "mfrac": [
{ "mrow": [
{ "mo": "-" },
{ "mi": "b" },
{ "mo": "±" },
{ "sqrt": {
{ "mrow": [
{"mi": "b"},
{"msup": { "mi": 2 }},
{"mo": "-"},
{"mi", "4ac"}
]}
}}
},
{ "mrow": [
{"mi": "2a"}
]}
]}
}
It's not as general, but works if you know your syntax is similarly bounded. I don't know how certain static languages would handle serial/deserializing, but makes construction via javascript literals much more pleasant. (mrow (mi x)
(mo =)
(mfrac (mrow (mo -) (mi b) (mo ±)
(sqrt (mrow (mi b) (msup (mi 2)) (mo -) (mi 4ac))))
(mrow (mi 2a))))
And this would be a saner one, where mrow is implied: ((mi x) (mo =) (mfrac ((mo -) (mi b) (mo ±)
(sqrt (mi b) (msup (mi 2)) (mo -) (mi 4ac)))
(mi 2a)))
I think either of those is clearly and inarguably superior.It has a single standard way of doing schemata that all the tools support, which is great. The maven pom.xml format is a much clearer way to specify a dependency than most of the alternatives (which often use an excessively clever concise form), and has really good autocomplete when editing it in eclipse (because eclipse understands the schema and so can offer autocomplete based on the elements that make sense at that point in the document).
If XML had just not bothered with namespaces I think it would have worked really well.
Could have been a bit simpler but hindsight...
XML, on the other hand, had the ability to do something using attributes in the element or a DTD / XSL. Occasionally I do miss that ability to communicate data schema alongside the data. But only occasionally.