Ask HN: Why did XML failed to be the standard for data format?
[1]: Yes, I know HTML is also a derivative in the family of XML/SGML, yet don't most of you frontend devs think that HTML sucks too?
[1]: Yes, I know HTML is also a derivative in the family of XML/SGML, yet don't most of you frontend devs think that HTML sucks too?
For data transfer now I'd actually prefer csv over xml due to the simple fact I can at least sanitise it. That is how low my opinion has gotten over years of data migration projects. In my new work I primarily work on JSON / csv / tsv and would not even consider xml. XML schemas have the idea of keeping data clean etc as well as guiding parsers etc but the added functionality comes at a price in performance and parsing complexity. I did find the xml transformation abilities useful however.
My personal opinions aside, friends of mine using java still swear by xml so if you're from that crowd your opinion might differ. I'm guessing the java infrastructure is more forgiving but its been too long for me to assume more than that. They obviously find utility in it: terabytes of xml is not trivial by any stretch of the imagination.
In its raw form XML is not much different from JSON. It is more verbose (e.g. start and end delimiters versus a name tag in JSON) but overall very similar, if not identical in terms of features. Schemas are a different topic.
Reading and writing with automated tools very much depends on these tools and not so much on the syntax itself.
JSON is slightly slimmer and less verbose but does not really offer any fundamental advantages. Both syntaxes are very similar. The reason why JSON is more popular nowadays can be most likely found its relatively close ties to JavaScript and JavaScript's success in web development.
- "Obviously conforms"? XML does too. Both are hierarchical document formats. Their syntax is different, thats about it.
- "Considerably slimmer"? That is arguable, I'd say "slightly". But yes, JSON is somewhat leaner. So what, that doesnt make it intrinsically superior.
- "Legibility"? That mostly depends on the formatting. Especially in today's world and its minifiers essentially nothing is legible any more. JSON, being less verbose, might have a fraction of an advantage here, but again thats about it.
Examples? Sure. Would you argue the following JSON document is legible?
[{"_id":"5e120086aa2b07d5af7ddda3","index":0,"guid":"ae823405-6305-4c29-a8d3-429423e0ff7c","isActive":false,"balance":"$2,616.37","picture":"http://placehold.it/32x32","age":36,"eyeColor":"green","name":"Lyons Pollard","gender":"male","company":"TECHMANIA","email":"lyonspollard@techmania.com","phone":"+1 (870) 441-2429","address":"722 Truxton Street, Osmond, Indiana, 1802","registered":"2018-03-14T10:56:23 -01:00","latitude":43.427473,"longitude":-78.16956,"tags":["id","voluptate","velit","sit","duis","velit","proident"],"friends":[{"id":0,"name":"Elba Fernandez"},{"id":1,"name":"Rosalinda Morrow"},{"id":2,"name":"Hannah Leblanc"}],"favoriteFruit":"apple"}]
Didnt think so. In comparison its XML equivalent is pretty legible <?xml version="1.0" encoding="UTF-8"?>
<root>
<element>
<_id>5e120086aa2b07d5af7ddda3</_id>
<address>722 Truxton Street, Osmond, Indiana, 1802</address>
<age>36</age>
<balance>$2,616.37</balance>
<company>TECHMANIA</company>
<email>lyonspollard@techmania.com</email>
<eyeColor>green</eyeColor>
<favoriteFruit>apple</favoriteFruit>
<friends>
<friend>
<id>0</id>
<name>Elba Fernandez</name>
</friend>
<friend>
<id>1</id>
<name>Rosalinda Morrow</name>
</friend>
<friend>
<id>2</id>
<name>Hannah Leblanc</name>
</friend>
</friends>
<gender>male</gender>
<guid>ae823405-6305-4c29-a8d3-429423e0ff7c</guid>
<index>0</index>
<isActive>false</isActive>
<latitude>43.42747</latitude>
<longitude>-78.16956</longitude>
<name>Lyons Pollard</name>
<phone>+1 (870) 441-2429</phone>
<picture>http://placehold.it/32x32</picture>
<registered>2018-03-14T10:56:23 -01:00</registered>
<tags>
<tag>id</tag>
<tag>voluptate</tag>
<tag>velit</tag>
<tag>sit</tag>
<tag>duis</tag>
<tag>velit</tag>
<tag>proident</tag>
</tags>
</element>
</root>
Yes, more verbose - which I already addressed - but nonetheless self-explanatory.
One advantage of JSON? Additionally to strings its core syntax defines three additional value types, whereas in XML everything is a string.XML and JSON are so similar they could be considered siblings. JSON's popularity does not stem from being "a cleaner expression of what xml was trying to represent" because everybody using it made a careful evaluation and decided after long deliberation that JSON is "Occam's razor", but simply because it is the default choice in JavaScript and comes with native support, whereas XML support is pretty shaky in Vanilla JavaScript.
That is why - to adopt the same confident attitude ;)
Would I slightly favour JSON over XML these days? Yes, probably slightly, but certainly not because it was better or offered things XML didnt.
JSON exists between the bureaucratic XML at the top and the anarchy of csv/tsv at the bottom. I'd rather operate on JSON data: fast to process, easily pulled apart, etc. CSVs are a free-for-all that are either well formed or a subtle spaghetti of loose commas and bad quoting.
What XML has that the others don't is verbosity and that helps in rebuilding damaged data. I can rebuild an XML record easier than JSON and XML lets me see where records are incomplete/malformed. These are its strengths.
But every new project I work on has JSON because the parsers run fast and everything is legible. I can edit a JSON file without it becoming a sea of text like XML.
To each their own, XML has a place and so does JSON. I know as fact that for my purposes XML uses much more time and memory and both of those cost money.
Yes, XML libraries tend to be unnecessarily complex (particularly with Java) and are often a drag, but this is not about about the format. Even the bureaucracy you mentioned is not inherent to XML itself. You can perfectly use XML without schema insanity, just like JSON.
XML is quite more verbose (which might help in a rebuild you mentioned) but is generally quite alike to JSON. Even the performance and memory issues you mentioned are not specific to XML but rather to libraries which went bananas.
My point is JSON is not more popular these days because it is inherently superior. JSON is more popular because it was the default choice in a JavaScript environment and - probably - because those who implemented JSON libraries were a tad more sane than their XML counterparts.
I find XML tooling far more bulky than that for JSON.
For data interchange between C++, python and lua plenty of support for JSON makes this clean and painless. The files are relatively clean and easy to maintain and fast to read/write. I'd not even suggest XML now since initial prototyping showed it too inconvenient for what we do.
As I said, XML on Java is often a hassle but that is neither XML's fault nor Java's but the responsibility of those developers who thought going bananas on interfaces, factories, and alike is actually a good idea.
As for XML, its syntax is somewhat more verbose than JSON's but thats about it.
-----
Just to provide one very basic example
XML
<people>
<person>
<name>John Doe</name>
<dob>2000-01-01</dob>
</person>
<person>
<name>Peter Smith</name>
<dob>2001-01-01</dob>
</person>
</people>
JSON [
{
"name": "John Doe",
"dob": "2000-01-01"
},
{
"name": "Peter Smith",
"dob": "2001-01-01"
}
]
Apart from XML being more verbose (but also providing more context), these two documents are virtually identical, they even have the same number of lines.And if one used XML actually idiomatically the example would slim down by a lot
<people>
<person name="John Doe" dob="2000-01-01" />
<person name="Peter Smith" dob="2001-01-01" />
</people>Parser complexity creeps in on that last example in particular. The first example was probably in the form the (possibly) easiest of all three examples to parse.
This all shows XML needing to support at least two different but related mechanisms. I'm not saying any of this is bad or that it makes parsing impossible: its just a source of complexity. JSON has key value syntax as well. XML has this flexibility but it quickly adds up. This is not a fault - it's the whole point of XML.
I'm not sure what your goal here is. I've got hard data for the work I'm doing showing XML doesn't provide significant benefits in exchange for its added complexity. Others likely have done similar for their purposes.
Matching the solution to the problem is generally a good idea. Are you perhaps too biased towards a single solution without considering the alternatives? I don't/can't know. Something to consider.
(I don't consider XML as useful now because it has shown itself no longer a match for the types of problems I'm solving YMMV)
I am afraid I cannot follow your argument about complexity as any parser worth its money WILL BE ABLE to parse XML without any significant overhead and - as I mentioned numerous times - issues do not stem from how an XML document is built but from the fact that (mostly) Java parsers attempted to implement about each possible OOP pattern instead of going for KISS. The code over-engineering is the issue, not the document structure. I am really not sure what point you are trying to make and why you are diverting from the actual topic I addressed.
I am slightly surprised about your bias remark, as it seems to be you who is strongly biased towards JSON despite my initial comments as well as the "significant benefits" as I never made any such claim, on the contrary it is you once again who seems to make such about JSON. Maybe you can post the "hard data" you referred to.
Bottom line (once again), XML and JSON are extremely similar and saying one is better than the other simply shows lack of experience. Then of course, if you parse 70 kilobyte JSON documents with a lean parser, but parse 12 megabyte XML documents with a typical Java parser, nobody needs to be surprised the latter will perform abysmally compared to the former, but that would be a whole different subject.
I'm not an evangelist for JSON, I'm someone who ran tests and came to conclusions with the help of multiple others. These weren't a generic benchmark for random or general academic purposes. These were representative samples of datasets we're actively going to be or actually are already using.
Even in my other interests I'm using JSON for configuration and data transfer. It shines there quite nicely. XML was generally suitable but its verbosity didn't provide any real advantage and the library support tried to drag in too many dependencies. TSV files weren't suitable even though they were simpler and we had control of the data sources.
You mention Java / Javascript but neither is what we're using. There's probably some irony in not using javascript for JSON i/o but it is what it is. (The purists will agree there's no requirement and so do we). You also didn't mention in passing any of the other interchange / file and document formats we actively compared. JSON / XML etc were just some of the candidates.
Thank you for letting me know that the teams I work with demonstrate "lack of experience". I've forgotten which logical fallacy that is but I'll leave that to someone else to know or look up. I'm just glad I've kept beginner's mind: its a key aspect of neuro-plastic mindset. Its a requirement for keeping an open mind.
We won't be ignoring our testing on real data subsets. The results are clear enough.
(The samples we tested with were around 5MB, 50Mb, 1GB, 10GB and 50GB in size as various combinations of lists and trees etc.)
However you havent really addressed your use case anyhow but you just threw keywords around - gigabyte, IO, compressed, etc. You might want to elaborate on where you have to use XML files of the size of 50 gigabytes.
It doesnt really matter what you are using, Java was just one example. If the XML parser you employ has similar issues you must not be surprised if the outcome is similar. And you seem to be coming back over and over to software support (dependencies). Yes, particularly Java was poor when it came to that but as I said quite some time ago, that is an issue with that software not the document format.
I will disregard the 10M+ tests, but could you publish somewhere the results of the 5MB files?
Again, JSON and XML are way too similar to be anywhere close to what you described and aforementioned benchmark highlighted that. Yes, its dataset is average but I am sure you'll be able to extrapolate that for larger sets.
Apart from the apparent improper use for data of that magnitude, I could only imagine you used an XML parser that simply was not fit for the task and if you do that you shouldnt be surprised that it does not work.
For the benefit of the probably only two others in the studio audience (who are probably currently both facepalming), we tested with multiple libraries, multiple languages, multiple OS and multiple data subsets. We found in our particular experience that JSON worked the best across our criteria using a representative sample of our datasets. Nowhere did I say XML is always the wrong choice for others. I vaguely recall I wrote I'm now none too keen on XML but have used it in the past. For some things I'd actually choose TSV over XML but thats on fairly, hopefully, obvious cases. I think XML's verbosity is actually its strength but that it has tradeoffs which are quite real. This should not come as a revelation to anyone.
I shared a necessarily limited snapshot of an experience I had and an opinion I formed based on it. I think others can do their own testing as I expect they will anyway. They will confirm or deny based on what they are doing. Especially the opposite case of large imports in XML being faster than everything else. That's completely fine by me.
You've definitely made too many assumptions based on too little data. You didn't even ask what industry this was for. Or what kind of data it was. Or even what disparate systems were involved such that we'd end up with something you state are inappropriately large compressed text files. You disregarded the use of "keywords" such as gigabytes or even compression in general as if those should be unimportant to us. Or why we would use JSON at all. Then you make judgements. Fairly condescending ones at that. This shows a general lack of awareness across several aspects of life in general. For the sake of both of those other people still following this chain, I'll finish here. Life is too short.
> However you havent really addressed your use case anyhow but you just threw keywords around - gigabyte, IO, compressed, etc. You might want to elaborate on where you have to use XML files of the size of 50 gigabytes.
I even asked if you could provide that one 5 megabyte file. I take your response as you cant.
I really have the feeling we are going in circles here and you seem to want to resort to ridicule at this point, which will make the discussion pointless.
I believe I have made my point very clear from the start, elaborated more than once what my stance on this subject is, and even dug out some benchmarks. If none of that pleases you or makes you understand what I was actually trying to say, then I am terribly sorry but it is pointless.
And I'd appreciate if you could point out where I was "condescending", as I would object to that, except for the "lack of experience" and I still stand by that given the information you have revealed so far.
http://www.navioo.com/ajax/examples/json/test.php
In the first example XML actually is a tad faster, in the second example it is practically a tie (JSON wins by three or four milliseconds).
[1] Are we seriously arguing about milliseconds when it comes to document parsing?
> This all shows XML needing to support at least two different but related mechanisms. I'm not saying any of this is bad or that it makes parsing impossible: its just a source of complexity. JSON has key value syntax as well. XML has this flexibility but it quickly adds up. This is not a fault - it's the whole point of XML.
It is not really two different mechanisms. It simply is part of XML's syntax, just in the same way as these two JSON examples are very similar
[
{ "hello": "world"},
{ "hello": "world"},
{ "hello": "world"},
{ "hello": "world"}
]
{
"1": { "hello": "world"},
"2": { "hello": "world"},
"3": { "hello": "world"},
"4": { "hello": "world"}
}
Does any of that add complexity? Yes of course it does. If you dont want complexity you should probably go with a "Hello world" program, which incidentally will be also relatively bug-free ;)However saying the additional syntax support of XML adds complexity which adds hours of CPU time when processing the same data in different formats (JSON and XML) borders ludicrous. Unless your XML parser is (deliberately) slow of course.
json avoids these pitfalls and people mostly just use a lib to generate it from native data structures, so it's almost always well-formed.