Deprecating XML
norman.walsh.name
norman.walsh.name
We were once, no lie, forced into JSON-over-XML because of technical and organizational imperatives that had to be reconciled with the need to ship functional code sometime this year. ("The web services parser is breaking again!? Screw it. We only need three fields. JSON grab bag, serialize as string, read string and interpret as JSON on other side.")
This is absolutely, positively, NOT XML's core use case. This is yet another case of Big Freaking Java Web Application developers finding a way to add gratuitous complexity, along with their dependency injected, aspect oriented, AbstractDataSourceFactoryFactoryImpls.
By the way, a quick Google brought up this, which is evidently not a parody but actual code:
http://www.docjar.org/docs/api/org/outerj/pollo/xmleditor/Fa...
The mixed content case does not occur in the vast majority of Big Freaking Java Web Applications.
One of these things is that you want to be able to pass objects over the wire without having to reduplicate all the code on both sides. For example, if you have Student objects in the student management system and also have Student objects in the university SNS, and those systems need to periodically pass data between each other, somebody might say "It would be great if they could do that without us having to custom-write data massaging routines. After all, we already have student code on both systems." and another guy might say "True, true, but we don't want to keep the Student class in sync on both systems neessarily." and then a choice gets made and someone forgets that one of the Students contains a graduation date implemented by a long-forgotten contractor as a wrapper class with some utility functionality which contains a Locale which cannot be serialized and then all of a sudden you are in the sixth layer of XML Hell.
Not that I'm arguing with you -- God willing, I will never, ever have to work with Java architecture astronomy again.
Really, HTML is needlessly verbose and annoying - we're just used to it. If you take a look at something like HAML (not that I agree with all of their choices), you can see how much nicer it could be.
That "needless" verbosity makes error correction much more robust. Haml has a shit fit over the slightest of irregularities. The XHTML experience should be proof enough that trying to enforce picky parsing rules or highly context-dependent syntax leads to breakage.
What? No, it doesn't. XML has even more structure than JSON does. Unless you're using it improperly, and the you're negating all the advantages that XML was supposed to bring.
It wasn't for 'mixed content'. It was for system interoperability.
The advantage that XML brings over JSON is that you can have an external document that you can use to validate any XML you're sending or receiving. You can guarantee it's formatted correctly. JSON provides no such guarantee.
Not that that's necessarily a bad thing.
Also XML brings other useful features like data types.
You can convert, losslessly, between XML and JSON. Doesn't that mean they have identical capabilities, though one may be easier in certain circumstances? There's no "need" in any one which the other doesn't also need.
IBM's, on the other hand, has a number of nonsensical statements like this:
>Because JSON is incapable of preserving any notion of a Base URI, it is unlikely that an application using the JSON serialization is capable of properly rendering markup containing relative URI paths. It is therefore important that such paths be resolved automatically during the conversion process as shown ...
Bull. Base URI is just a field, or inferred by the document's source. ie: the same as XML. Their technique is incapable of translating a Base URI, it has nothing to do with JSON. Theirs also cannot handle namespacing, while Google's is simple, and almost identical to how XML does it: namespace$field: value.
A lot of these conventions are available the XML standards. Of course if they had started with something like json the result would be less verbose but maybe not that simple either.
There's no reason a JSON format can't do the exact same thing. JSON-RPC, for instance, defines its own rules on top of JSON which anything using it must conform to. How hard would it be to make a set of JSON interchange rules which specify external schema documents?
edit: answer:
{json_strict_version:1, schema:{external:[url,url], internal:[{schema},{schema}]}}>Conforming XML processors fall into two classes: validating and non-validating.
Yup, that's a guarantee alright. A guarantee consisting of the people who wrote the parser agreeing to write DTD-parsing code. Or not. Depends on your parser.
It has a specification for validation because a specification for validating it has been written, based on agreement of the specification and implementation by the parsers, not because XML has some inherent, magical validatability quality.
edit: this is all on top of that a spec is just a spec until someone implements it. Plenty of specs have been chucked because others did something else, or have attributes which nobody implements (how many email forms accept the full RFC-compliant email set?). A spec is an abstract agreement.
And then we're 10 years into the future and we're seeing this exact same article again, except with s/JSON/some simple JSON replacement/g and we're back to square one.
JSON has its uses, it's quite nifty in its ease of use for doing ajax stuff, but imo the point of the article was to point out that xml vs json is a nonsense comparison much like 'sql' vs 'nosql' databases.
Different XSD implementations do have slightly different ideas on interpreting the schema. (e.g., if your data has a bunch of A tags, then a B tag, then another bunch of A tags, do you include two occurs:unlimited statements for the A's or just one?)
By the way, Javascript works just fine as a "query" language for JSON. Not as good as XPath or XQuery work for XML, but definitely better than any XML DOM manipulation in Java or another language would.
Now the standards can totally suck - they can be improperly implemented - and they can be difficult to make performant... But those are largely issues of implementation unrelated to the core XML and XML namespace specifications.
It seems unlikely to me that JSON could reasonably be purposed to those use cases...
Since any XML document can in theory be converted into a complete JSON representation I do believe that nothing is keeping people from doing exactly those use cases. You're right in asserting that this isn't happening in practice with JSON. I allege the reason for this is that XML isn't really used for these kinds of operations either. Sure, the XML working group put out a stunning amount of "standards" that prescribe these conventions, but in practice NONE of it actually works across vendors without heavy intervention and some pretty shocking compromises in application code. Most complex XML products are not actually that interoperable, they mostly just work because one or more communication endpoints are basically faking compliance of their interfaces.
How? (apart from a char array...)
# <a href="http://example.com">hello world</a>
{
"node-type":"element",
"namespace":"http://www.w3.org/1999/xhtml",
"name":"a",
"attributes":[
{
"namespace":null,
"name":"href",
"value":"http://example.com"
}
],
"children":[
{
"node-type":"text",
"value":"hello world"
}
]
}http://seanmcgrath.blogspot.com/2007/01/mixed-content-trying...
It's not as persuasive as it could be, but it gets the idea across.
And in any case, in order to handle truly inline content you've still got to correctly detect such in-line code, and label it as separate from regular text which just looks like inline code. Which means escaping. Which means
"text {b:'bold text'} more text"
is the same solution, with the same problems.>Here is XML's sweet spot (using square brackets to keep everything un-mungable by the angle-bracket chewers in my current tool chain):
Sweet spot, in that parsers usually just assume everything which looks like code is code, even when the spec allows for other ways, forcing people to fake it with things like square brackets.
edit: whups. Editors, not parsers, are primarily responsible for such mangling, by making it hard to insert code-as-text. The problem with inline markup still exists, however.
I hate XML because it is so mis-used. People try to use it for EVERYTHING and you get these monstrous XML trees with legions of attributes, when what you're trying to send could comfortably expressed in "{success: true, message: 'Record deleted.'}".
I admit my ignorance, at least. I'm sure that were I properly shown well-used XML and these wonderful and vast tools that the competent XML worker has at their disposal, I'd be as happy to work with XML as I am to work with PHP.
That said, you won't ever catch me trying to write a 3D FPS in PHP, and I hope that if you ever DID catch me doing so, you'd give me the exact same fish-slapping that XML abusers deserve.
As in, XML nodes are "game script commands" which are executed sequentially.
I believe you have come up with a rare example where XML is decidedly better suited than JSON!
However, the majority of XML use cases are for document-interchanging applications and the way those are being handled requires massively bloated toolchains of poorly inter-operating software.
> If you want a language with a JSON-like syntax, consider using JavaScript.
You're probably right in making this joke though, since JSON is based on JavaScript it certainly owes its popularity mostly to the browser environment.
I don't think that's a given. JSON is useful because it provides a clear, compact syntax for the most commonly useful data structures - lists, dicts, strings, numbers and booleans. That's certainly why I switched to it over XML (and serialised PHP objects).
Why not just... use a scripting language?
I hear LUI is pretty popular as a DSN these days.
But Protocol Buffers are best. Like JSON, the data model for Protocol Buffers maps nicely onto simple data structures (unlike DOM). But with Protocol Buffers you also get a schema for free, a wicked efficient binary format if you want it, default values, etc.
And you can use JSON as a text format. In other words, you could take your JSON that you have sitting around, whip up a Protocol Buffers .proto file for it, and get nice generated C++ classes for it with full schema validation. Your JSON file:
{field1="foo", field2=5}
...could be accessed from a C++ object as: my_obj.field1();
The only bummer about protocol buffers is that its support for high-level languages (PHP, Perl, Ruby, etc) is not very good. I've been working on a separate implementation of Protocol Buffers in C to address this (by making it easy to write bindings for) but this project has unfortunately been stalled as I've been busy with work and life. :((I wrote a generic command-line Protobuf-JSON converter).
The protobuf tradeoffs only make any sense if you think you can somehow do something useful with messages that are somewhat but not entirely corrupted, because you still have to solve the problem of finding intact field boundaries without being given any HDLC-style framing.
[citation needed]
One of the features of protocol buffers is that they're backwards and forwards compatible. You can add and remove fields, change "required" to "optional" and back again, and still make sense of what comes to you on the wire. I don't think there's much of anything that can be eliminated from the protobuf binary format.
Eliminating tags for a field (even if you consider it required) wouldn't be backward compatible with a previous version of the protocol that considered it optional.
A message from a previous version of the schema is not "corrupted." Being able to make wire-compatible changes to the protocol is an extremely important feature.
Just because a field says "optional" doesn't mean it's logically optional. You don't have to make your schema formalism complex enough that it can describe every last rule of what it takes for a message to be valid. In fact you definitely don't want to do that, because it's a horrible amount of complexity in the schema for little gain.
Yes, it's true that some Googlers use "optional" instead of "required" everywhere in their .proto files. That doesn't mean that you can omit any field and expect your message to be processed by your peer without error. It just means that you won't get an error at the schema validation level. But the application could still throw an error. More complex rules about what fields must be specified or what values they must have can be described in comments, and enforced with custom validation if necessary.
Also, since protobufs support default values, you can define what value will be returned for scalar fields if no value is explicitly sent. This can often be used to define useful default behavior for the case where a field is omitted.
Not true at all. The server can implement them -- once -- and any client who makes an invalid request to the server will get an error message. These constraints can be expressed in comments in the interface (.proto) file.
> Like HTML vs. XHTML—I shudder to think how much work was wasted trying to handle the worst tag soup imaginable, simply for lack of a well-formedness (or DTD validity) requirement.
HTML and XHTML is a completely different ball of wax. Insisting on even well-formedness is simply unreasonable in practice, because it is so difficult to ensure, and it is the user who pays the price when the software isn't perfect.
If you're still convinced that the world would have been better if strict XHTML had won, you should read: http://diveintomark.org/archives/2004/01/14/thought_experime...
And as for compression: dfsch's binary serialization format does some de-facto dictionary compression on some fields and it actually speeds up decoding itself (decoder caches various high-level metadata in decompressor's dictionaries), not only reduces I/O size.
EDIT: also the concise yet all-encompassing nature of its home at http://www.json.org
I do have to agree with the author though. This feels like a really "meh" discussions.
Unfortunately, neither use case is typical. Usually it's a mix of both, and it's not worth using both formats so you have to pick just one.
But generally what I want out of a data stream is a set of name/value pairs. That's it. I really don't understand why, but trying to get those name/value pairs from an xml based api is always an adventure. Well - I know why. It's because everyone structures their data differently when they create xml. Sometimes the name I'm after will be a node name. Sometimes it will be a value of a text node. Sometime the value that properly maps to that name will be a textnode 3 layers deep, sometimes it will be an attribute value.
I'm still a newb but I've worked with about 8 different apis now. For the json apis I wrote one recursive function that worked on all of them to get the data I wanted... It was about 8 lines of code. For the XML apis, I STILL haven't figured out how to write a recursive function which works for ONE of them, let alone all of them. (while relying on python minidom to do the actual parsing for me).
Then SOAP happened. The same philosophy as XML-RPC gone horribly wrong: a human readable format no longer readable by humans, insanely complex, and hard to implement properly (even with the libraries). Frustrating all around.
Although JSON is much simpler than XML, there's no reason why someone couldn't invent something as horrible as SOAP (or XSLT) on top of it. I can only hope that JSON implementers see the value of keeping things simple. Perhaps the existence of XML as an "enterprise" technology will help differentiate JSON in that way.
var foo1 = new Foo1();
foo1.bar1 = new Bar1();
var foo2 = new Foo2();
foo2.bar2 = new Bar2();
If you were to serialize that data using XML you might do this: <Foo1 name="foo1">
<Bar1 name="bar1"/>
</Foo1>
<Foo2 name="foo2">
<Bar2 name="bar2"/>
</Foo2>
But now what if you use JSON? You'll probably end up with something like this: {
"foo1" : {
"class" : "Foo1",
"@bar1" : {"class":"Bar1"}
},
"foo2" : {
"class" : "Foo2",
"@bar2" : {"class":"Bar2"}
}
}
Generally it makes sense to use XML for imperative data structures, and use JSON for functional data structures. <Foo1 name="foo1">
<Bar1 name="bar1"/>
</Foo1>
<Foo2 name="foo2">
<Bar2 name="bar2"/>
</Foo2>
[["Foo1", {"name" : "foo1"}, [
["Bar1", {"name" : "bar1" }]
],
["Foo2", {"name" : "foo2"}, [
["Bar2", {"name" : "bar2" }]
]]
I personally wouldn't do this, but the arbitrary encodings between the two samples are now equivalent.Plus, serialization to JSON in a dynamically-typed language is pretty much automagic compared to what you have to do with XML.
In other words, I've seen many special cases of the mistaken belief that programmers of today can do things that only researchers of the future wielding human-level AI will be able to do -- and that just has to cause grief if those hopes and expectations inform IT decisions.
To generalize my viewpoint: XML is a good way to mix data and data processing, or mash up multiple DSLs in polyglot fashion. When stream-parsed, one can imagine great use-cases of XML, where an early node adds context, references and functionality for other data later in the document. But it sucks for plain old standalone-structured data, which is the kind of data we usually like to _store_.
But to deprecate XML, I don't think so. Imagine just an EDIFACT document in JSON?
Quote: "any damn fool could write a better data interchange format than XML"
json lacks that level of support from mainstream relational databases, which makes it a waste of time for certain use-cases.
I just like to point out that there was a time before XML. Fixed width formats are still very common when working with mainframe systems. Surprisingly, they aren't actually that bad to work with.
I still love JSON, http://ilovejson.com for it's readability and simplicity.