XML Can Give the Same Performance as JSON
infoq.com
infoq.com
XML is a poor choice for most serialization since it has the wrong information model. It's designed for extensible document markup and most application data packages aren’t documents that require markup.
The correct way to load an XML document is as follows, IMO:
* Open it
* Run it through the schema to validate it's
structure
* Manipulate/access it via XPath, with the ability to
assume correct structure as defined by the schema.
Compare to JSON, where the process is * Open it
* Attempt to access it, even though it's structure
may deviate from the expected
* Die or try to rollback work when an assumption about
the data structure isn't met.
Thus, most JSON code either trusts the input data implicitly, or is a mess of "is this where it should be?"/independently re-developed schema equivalent every time. I am aware of attempts to create JSON schema languages - unfortunately these are neither in common use, nor well defined enough to be standards, and they're more verbose and less expressive than XML's RelaxNG compact syntax.JSON is fundamentally a dump of a data structure, not a document format. It isn't designed for long term data storage, or for inter-compatibility, despite being used those ways. Thus, it's great if you control both ends and your data structures don't change, as is the case with most web programming. It frequently falls down outside that realm.
Now, the above ignores runtime performance concerns. I'd argue (again) that outside of web apps, the load/store into a data structure is done rarely, or can be done in the background or in parallel if there's a lot of data concerned. Also, we're probably talking about less than an order of magnitude in the vast majority of cases.
The advantage of a schema is that one schema can be used in multiple languages. Someone's Lua based embedded system can run the same schema as a Go program on a 32 CPU box.
You can give someone a well written schema, tell them "make your data like this, and run the validator when you open and save your XML". In this way, it's a free, one line (in most languages) structure sanity check.
In a typed language the very first stage of deserialization, the "turn this from a string into objects in my language" step, will do that for you "for free".
In an untyped language I guess it could be helpful, but again: it's almost always easier to write these sort of checks in a general purpose programming language than in XML schema.
>The advantage of a schema is that one schema can be used in multiple languages. Someone's Lua based embedded system can run the same schema as a Go program on a 32 CPU box.
If that's what you want there are more robust and performant approaches, e.g. thrift.
>You can give someone a well written schema, tell them "make your data like this, and run the validator when you open and save your XML". In this way, it's a free, one line (in most languages) structure sanity check.
Right, but that doesn't tell them what they think it means. It doesn't make the request valid.
> JSON is fundamentally a dump of a data structure, not a document format.
Exactly! IMHO, it's unfortunate that mark-up documents can be used to represent other data structures. It isn't what they're meant for, and there are (or "should be" if you're not fond of JSON) better tools for human-readable serialization.
The original name (Yet Another Markup Language) threw me off, so I never seriously considered using it. Even its Wikipedia article is in the "Markup languages" category.
Yaml for configuration files/things people touch
XML for documents that are complex/nested, that need extra validation, or need to be stored for longer than a few days
JSON for ephemeral data structures passed between two controlled endpoints
SAX style streaming/pull parsers tend to deal better with this kind of thing as well.
It took years and lots of committee meetings for a W3 blessed XML schema language; and then, key people didn't agree with how it should be done: James Clark's alternative was TREX and MURATA Makoto's was RELAX.
Every significant project I've worked on has the same cautious loading process for XML as JSON, with some extra checking at the early XML load stage because XML is harder to work with and thus fewer people produce valid documents (forget schema validations, errors with simple character encoding, well-formedness or namespace declarations are surprisingly common). In practice, I tend to end up with a forgiving parser, a collection of selectors and full validation on the results, which works equally well with either format.
This is, and always has been, an absolutely terrible way of writing off XML. Do we judge every technology today based upon their purported original purpose? "The computer was designed to solve linear equations, therefore it is ill suited for any other purpose". Noisy nonsense.
JSON only maps to the information model of JavaScript, and has no closer ties to any other language than XML does.
By contrast, XML does not directly map to common data structures.
If you believe that XML's information model maps onto "any other language" just as well as JSON, please show me native mapping/object support in XML and, conversely which programming languages (besides XSLT) have mixed content and attributes.
You can serialize to/from XML in virtually any modern language. Not sure if you're seriously asking this, or what the limiter "native" is supposed to mean (beyond that JSON is JavaScript. But it isn't C#, and it isn't Python, and it isn't...).
How they use attributes or elements is context specific. Every feature and pattern doesn't apply to every use.
XML specification is ten times longer; it has dozens of sharp edges you have to memorize in order to use effectively.
No, you actually don't. You create an XSD of a simple layout (trivial, monkey work). That is what you serialize from/to. The universe of possibilities in XML is, again, utterly and absolutely irrelevant.
Seriously, JSON doesn't even have a DATE type. That is so fundamentally broken I don't even know where to start (Microsoft has their own mystical blend of date when they use JSON, for instance, incompatible with anything else).
JSON is used because Javascript, Python and Ruby (among others) understand it without needed external libraries, and because it's fairly readable for humans.
Performance has nothing to do with it. If you want performance and have some control on your stack, you're better off with something more specialized.
I don't think neither this is the point: afaik browsers (where the most of the js is evaluated), python and ruby also have xml parsers in the "standard" lib (maybe more than one).
What really matters is that javascript, python and ruby programmers undersand JSON without crying for the "<" and ">" presents in the xmls
Consider for example trying to map:
<some-data>
<a>...</a>
<a>...</a>
<b>...</b>
<a>...</a>
</some-data>
You can automatically map the first sequence of a's to a "aList" collection in your programming language, but then what is the automatic mapping of the one or more a's that follow the b element? Should it be a list or a single item ? Or is it semantics preserving to just combine it with the previous collection? If it should be mapped to a collection, what would that collection's name be in the target language, so that it doesn't clash with the previous collection's name? Etc. Formats like Json avoid this problem by only supporting sequences as explicitly named collections, like most programming languages do within objects or structures.IOW with XML you can't really do a proper mapping without an accompanying schema or other metadata, which is not true for the Json data.
If you have an a set of people, do you separate them out into a "manList" and a "womanList", and then panic about having to create a "secondManList" because some men came in after the women? If order is only important within gender, then you can happily just add the men onto the end of the original manList. If it is important across all people, then you have to have a "personList".
Just like you might have a "personList" that can contain "male people" and "female people", you can automatically map the whole sequence of "a elements" and "b elements" to an "elementList".
In the case of mixed content (text nodes and element nodes together), you would simply map them all to a "nodeList".
Also, do you really expect an automated process to do the abstraction from "female" list and "male" list to person list? That's more than just metadata that you're assuming to be in context (which was my point), but actual AI! This is very different from programming languages, where collections of things which occur within objects and structs are given names which reflect their intended meanings, like with Json.
How is this different from a Java object which contains two list members called "attributes" and "children"?
How is this different from an S-Expression containing two lists?
How is this different from a json object like {"attributes":[...], "children":[...]}? Bear in mind that there is no requirement for JSON lists to be homgeneous. {"things":[1,true,"hello",3,{"addressee":"world"},[{"greeting":"Hola"},7],false]} is a perfectly valid JSON object. You don't have to define it as {"numbers":[1,3], "strings":["hello"] ...}.
It is a pretty common behaviour, that if you want a homogeneous list of things that differ, then you make abstractions until the differences disappear, e.g in an OO situation, you go up the inheritance tree until you are at the lowest common base class. In an XML situation, that common base class is "Node".
Even in a strongly typed language that requires homogeneity in lists, the only thing you know about the list members is that they can be cast to the same type, not that the members are of that type and no other, and certainly not that they all have the same name. Consider a C++ array of CFruit objects. It may have a member that is of class CBanana, one that of class CApple, and another of class COrange. If you want to do something COrange specific with the oranges, then you have to perform dynamic_cast<COrange> on any member of that list you suspect of being an orange. The same is true in a duck-typing situation.
The reason for my male & female example is that you would normally simply have a list of people. Of course, if your model contains no base that is common to both men and women, then you can't expect the machine to work that out, but if Man and Woman both inherit from Person, or if there is no Man or Woman, just Person with a member that specifies a gender, then it's trivial.
Way too glib. There's lots of good reasons you may choose to go with that, and there's lots of reasons why you may not. DOM-based approaches pay a lot of resources for their functionality. If you don't use all of that functionality, you may lose over a SAX-based approach. If you're resource rich, hey, great, go with DOM unconditionally, but not all web apps are in that situation.
It's one advantage the XML ecosystem has over the JSON ecosystem; while JSON can be parsed in a streaming manner, you can generally count on a good XML event-based library for your environment, whereas streaming JSON parsing libraries seem more unusual. If you need it, you can count on it in XML.
JSON has been in Python's standard library since 2008.
http://docs.python.org/2/library/json.html
It's also in the Ruby standard library.
http://www.ruby-doc.org/stdlib-1.9.3/libdoc/json/rdoc/JSON.h...
JSON has been launched around 2002.
Python has incorporated one of the many existing libraries in 2008 after a lengthy discussion, the last part of which is http://mail.python.org/pipermail/python-3000/2008-March/0125...
Ruby had no support JSON for years, it has been added, in its current form, to the stdlib only in 2007.
XML has been supported by Python since 2000 (minidom + sax) and by Ruby since 2003 (REXML).
The idea that JSON is used more than XML just because programming language supports it "natively" clashes with the evidence.
(web services - JSON; configuration files - JSON for simpler files, XML for more complicated ones AND only if JSON couldn't handle it).
Also, I might be missing something but it appears that this benchmark doesn't take into consideration libraries used server-side to serialize and deserialize JSON / XML? I'd say that depending on those, there can be some big differences.
Not mentioning the fact that bigger impact on the feel how "snappy" application is lies in the overall application design, than in choosing whether we are using JSON / XML / YAML / anything else to transmit the data.
For me, a disadvantage in JSON for configuration files is the lack of comments. My current project uses JSON for config and I sorely miss 1) comments and 2) unquoted config var names. I'd rather use YAML or another format with a config loader for whichever language I'm using.
python -m timeit -r 10 -n 100000 -s 'import json' 'j="""[{"t":"Hello"}]"""; json.loads(j)[0]["t"]'
100000 loops, best of 10: 9.07 usec per loop
python -m timeit -r 10 -n 100000 -s 'from xml.dom import minidom' 'xml = "<t>Hello</t>"; minidom.parseString(xml).getElementsByTagName("t")'
100000 loops, best of 10: 74.9 usec per loop
python -m timeit -r 10 -n 100000 -s 'from xml.dom import minidom' 'xml = "<t>Hello</t>"; minidom.parseString(xml).firstChild.firstChild.wholeText'
100000 loops, best of 10: 76.6 usec per loop
Given that the "study" comes from a "markup conference" that has XML all over its website and given the fact that XML documents (uncompressed, in memory) are much larger and that XML is a much much more complex standard, i seriously doubt the conclusion that XML and JSON are "almost the same" in terms of speed and memory usage.Or simply put: How is it technically even feasible that an XML document that, by its nature, is much more complex (and thus i imagine a XML parser to be much more complex) can be parsed as fast as JSON with more or less the same memory footprint?
Wouldn't that only suggest that the JSON parser that is tested is just not as optimized as the XML parser?
I can imagine that a webbrowsers XML parser is much more optimized and mature then its JSON parser, given that it's a browser that mainly needs to parse HTML/XML?
Does it mean i should switch to using XML instead of JSON in commandline tools?
Or to put it another way: The headline "XML Can Give the Same Performance as JSON" is probably true for browsers which have had years in improving the XML parser. But i don't think this can be a general conclusion.
So? The overwhelming majority of JSON advocacy comes from developers who happen to know JavaScript, and thus JSON appeals to them. Virtually everyone speaks from the position of self-interest.
Until someone replies with the same experimental rigor to refute the findingsthen I think the observations this paper makes stands ... thats how peer reviewed journals work.
The paper is not saying 'XML is better' or even 'XML is faster' ... its just addressing the perception that XML is slow in certain scenarios which has become a default myth.
JSON has been accepted as a datatype in the Markup conferences of the world its great for data transfer, XML is a compromise on many different levels but tends to be good for mixed content and documents. I think we've all moved on.
How does this advice stack up against the recent security issues with compressed HTTP traffic? Is this article's recommendation at the same place in the transmission stack where this would cause trouble?
It doesn't. The recent issues with HTTPS suggest there may be a fundamental tension there between performance and security. In fact the recent issues don't particularly care "where" in the HTTPS connection the compression occurs, it just has to be inside of it. It won't matter whether you use standard HTTP compression or roll your own (which on the TCP socket won't look all that much different anyhow, you'll just be giving up browser support for automatic decompression).
Whenever I see a modern API that uses XML is gnash my teeth and shake my fist skyward.
http://balisage.net/Proceedings/vol10/html/Lee01/BalisageVol...
So his two JSON test cases are eval and jQuery. He does not use JSON.parse.
So this shows that with a lot of hand waving around cases no-one likely cares about, XML is almost as performant as JSON, even if the code is way, way uglier.
In an actual real project where you're probably passing lots of fiddly objects around we have no results.
speed: use thrift or protocol buffers
ease of implemention: json
xml has best of both worlds. And therefore it is most of the time not suited. However the fact that it can support schema (XMLSchema), document translations (XSLT/XQuery) and query mechanisms (XPATH/XQuery) makes it a format very well suited for big enterprises. Where specification is important.
This is why cluncky protocols such as SOAP are built on top of XML. It is future proof (extensibility is more difficult in json). It is schema based (parsing, validation and language binding is easier).
JSON for example doesnt support references between nodes. You can built it in, but its not standard.
JSON is the easy peasy solution, the quick win, the fast enough one, and therefore the winner. However XML is the big beast that has it all, and therefore its complex. But that doesnt mean it sucks.
This title is misleading, nowhere does it have the word "browser".
Makes me really wonder why they went with that format.
The major advantage of JSON is a readable syntax. XML tends to be overly verbose, and therefore not as easy to read.
JSON is more compact too, which matters when sending 1000s of objects across the wire.
That's pretty subjective. For example, I find XML to be far more human readable than JSON.
Developers choose json over xml because using it and dealing with it is more lightweight than dealing with the very baroque xml. The fact that the typical json payload is 1/2 the size of xml is just a bonus.
CSV Can Give the Same Performance as JSON
for data transmitters: lean towards json unless your data structures are deeply nested.
for data receivers: be prepared to handle both (as even in cases where json clearly makes more sense) as large orgs tend to lean towards xml. When an API offers both serialization options, choose the better one so through log analysis you can nudge the producer towards the optimal solution.