XML is almost always misused
devever.net
devever.net
Whether that was the intended purpose when xml was designed is irrelevant. It’s what xml is used for in almost every case.
The author also doesn’t suggest what should be used instead to encode structured data, or perhaps more importantly what should have been used to encode graph like things such as map/lists/objects in the 2000’s. Json really hasn’t been an alternative until quite recently (10 years ago?).
In fact reading the article carefully I fail to see the author argue why xml shouldn’t be used as a data format either.
XML is a complete mess. Have you SEEN it's spec?
You can put JSONs spec in a single page. XMLs spec, not so much. Hell, most of the XML parsers don't support the spec, and the ones which do, historically have been riddled with security holes.
JSON over XML was simplicity over a crazy spec built by a bunch of companies all wanting to shove their own crazyness into it.
https://en.wikipedia.org/wiki/JSON#Data_portability_issues
XML has XSD and RELAX NG for more than 15 years now. https://json-schema.org/ is still a draft.
It’s not as simple as saying “everywhere xml is used, json would be a better choice”.
The fact that MobileDoc exists makes me physically ill. Something that can be expressed with one line containing a paragraph element and an italic tag is over a dozen lines of JSON spam.
Yes, storage is cheap. Bandwidth is not. Also, you really don't want a human that intercepts your data to be able to read it. Additionally, your data structue only makes sense in the context of your domain, which usually has been modeled in your program(s) that work in that domain, and thus it will be better if you deserialize it within tools that understand that domain.
If you feel the need for a general purpose deserialization protocol, there are several available - Avro/Protobuf, etc.
Binary encoded data can often be decoded without consuming the entire document. Sax-paparser-like reading requires at least reading an open and a close before the data is useful.
String serialization is a wasteful endeavor. It makes life easy for devs because it takes one less step to read the data in a text editor or log message, but quite often requires hacks to model things like recursive or self-referential data structures, and wastes space by repeating property names constantly for every item within the serialized structure. It's predicated on four of the fallacies of distributed computing, namely - bandwidth is infinite, the network is secure, transport cost is zero, and latency is zero. It is a solution looking for a problem, and because we are lazy, we don't build tools that would make binary serialized formats just as easy to use as json/yaml/xml.
I use csv when applicable. I use protobufs when applicable. But for the typical use case I choose xml for it's some config/dsl/dataset that needs to be human-editable (support comments, for example), more complex structure than csv supports, and preferably not need an external library or a custom parser. Json, Csv, Toml, S-Expressions, protobufs all fail one of more of these requirements. I'm sure there are others but none that don't have at least one drawback I don't want.
A poor man's data storage is exactly what I want!.
And, being honest, JSON schema is better than, say, GPB or Avro schema at enforcing field relationships, e.g., "if typeId is 7, then partnerId cannot be null"
And XML doesn't? Quite a few (not all, but quite a few nonetheless) programming languages include zero support for reading or writing XML-formatted data without using an external library or custom parser. This includes nearly all languages that predate XML, and quite a few languages that postdate it. Even when a language does have built-in (or at least in the standard library) support for XML, it's almost always a royal pain to use, especially once namespaces and schemas are involved.
Once upon a time, though, the answer was (and in a lot of places still is) INI:
- It's human-editable and supports comments
- It supports more complex structure than CSV
- Some languages have built-in support for it, and the Windows and GLib APIs support it, too (well, something similar enough to be compatible, in the latter case)
INI falls flat when you need to express deeper levels of nesting than keys and sections, though.
There's also YAML, which meets all your criteria about as well as XML does (at least on average; your specific language/platform might favor one or the other).
Omg YAML. Here's an example for you.
- external:
metric:
name: kafka_consumer_group_lag
selector:
matchLabels:
topic: rtb_trx_records
consumer_group: trx-record-validator
target:
type: Value
value: 30000
type: External
It seems like the "type: External" should line up with "metric" and "target" but no, it needs to line up with the word "external" - not the dash, but the word after a space after the dash. Using YAML frequently reminds me of the quote "Be open minded, but not so open minded that your brains fall out".On a platform that has almost no support out of the box (e.g python) the choice is open. But on a platform that has a couple of formats built in, picking a format outside that platform is a pretty big step. The return needs to be substantial for a .net developer to use yaml via an external library over xml.
My reasoning in this thread has always started from the perspective that xml comes built in and almost no other format does. This is the case for e.g java and .net but not for python or C for example. But the prevalence of xml comes from java/.net so if we are to ask why, then we should consider that.
Make an xml stylesheet and your kubernetes cluster is instantly documented.
https://www.gnu.org/software/emacs/manual/html_mono/nxml-mod...
I am writing non-XML code most of the day, and I do not have structured editing / auto braces enabled. So when I need to edit that one XML config, I'll open it in my regular editor, which will provide at most syntax highlight, and edit it as needed with a bit of swearing. And next time, I would promise myself I'd choose a different config format which does not need special editors.
I myself don't even use syntax highlighting and normally work in vim and although I do make errors in XML sometimes, I find that I make at least as many syntactic errors in Python or C code that I have to weed out before I can proceed. But I never heard anyone complaining about Python or C being too strict :)
That sounds like a very passive-aggressive way to deal with a problem. Do you do the same thing when writing programs?
In general, when you see something inefficient, you can either fix it to make it better, or ignore and come up with random workarounds.
In my opinion, a config file which cannot be edited by hand, and which needs a special editor with non-trivial learning curve, is a inefficiency. I can either ignore it, and set up the specialized tools; or I can fix it, by ripping out XML and replacing it with something more human-editable, like TOML or YAML. In large teams, it is almost always better to fix it -- sure, I will spend a few hours getting rid of XML, but this will pay itself off in the long term, as no one else will have to bother with special setup anymore.
(This obvious only applies to the systems where XML is a minor part, like a single configuration file. If your system has huge amount of XML, you better learn the right tools)
Not even formats designed for human consumption such as yaml are very good. The good ones for editing (toml, csv, ini) fall short when it comes to complex structure instead. There is no silver bullet.
It’s testable, discoverable etc.
(Sorry about shameless plug, I’m not affiliated)
Having to use a gazillion declarative languages to achieve what a regular programming language does is simply crazy.
<invoice id="123" customer-id="456" date="20199-10-30">
<item no="1" product-id="789" qty="42" />
</invoice>
Is this really a poor man's choice? And compared to JSON?! I can see at least the following advantages here:1. Each element has explicit type name (invoice, item). JSON is "typeless", which simply means the type information travels out of band. And with XML namespaces these type names can be made globally unique, but still stay human readable.
2. Each element is self-contained, the code that produces the <item> doesn't need to know if there was an item before or after it so that it should add a separator. (The dangling comma problem in JSON.)
3. The attribute names are not just arbitrary strings as in JSON, there are strict rules of what can be in the name. They're much more suited for structured data than JSON, where you can name an attribute "foo.bar" and some JSON readers that accept a JSON "path" won't be able to find it.
4. It has less visual noise than JSON because the attributes don't have quotes around them and you don't need to separate elements with a special symbol. Despite the common belief well-written XML is more readable than JSON.
And we haven't event touched things like validation + extended types, references, and transformation of data.
Lisp?
It may not be the popular answer, but it’s definitely a valid answer.
The advantage XML would have in that situation is that, because it's so verbose, if you get a malformed XML, you can eyeball-parse it and often figure out how to hand-edit it to make it valid. If it is valid, you can also see exactly how the schema you were sent differs from the schema you expected. An S-expression, having less redundancy, also is potentially more brittle.
It's a selling point while you're trying to figure out how to get it working. Once you have it working, it's not - but by then, you have it working, so why change it?
Let's not forget, many one of these configuration files and data formats are one-off hacks that were meant to be replaced by a real format, a real parser, a real DSL etc. The reason the xml config/dsl/format stuck is because it worked. And it was cheap and easy.
((foo . (bar . baz)) (foo . (bar . baz)) (foo . (bar . baz)) (foo . (bar . baz)))
And say "index all the foos", but you're mixing the structure and the content, in a way that JSON and XML explicitly separate.
> ASN.1 is a joint standard of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) and ISO/IEC, originally defined in 1984...
Today, of course, we treat the entirety of our deployed infrastructure as 'merely' a platform to write code. And not only are experimentation and failure OK, they're positively encouraged. Velocity became important.
Then, once you deserialize it, it's still a printable version of ASN.1. Sure, it's unambiguous, rigidly defined, and standardized. It's still gouge-your-eyeballs-out horrible to try to do anything with.
Say you get an XML message over the wire with a bit flipped. If you look at it, you have a good chance to be able to figure out what went wrong, edit one character, and you can now process it. If you get an ASN.1 message in the same condition, it's pretty much game over (though there may be special tools that could save you).
Say you get an XML, and you don't know the schema. You look at it, and you can see what's going on. You get an ASN.1 where you don't know the schema, and you can be totally sunk. (If I recall correctly, in ASN.1, you can have schemas that are private, that is, not specified in the standard.)
XML was vastly overused for a long time. That doesn't make those usages correct, as there were alternatives even then. (It also doesn't make people who overused somehow bad; I think it was a reasonable and necessary mistake.) It certainly doesn't make new ones correct now given that JSON's been around 16+ years. [1]
I think the author here is slightly strong in his criticism; I think XML is great for things that are meant to be long-lived and self-documenting. That is, things that are used like documents. But if I'm passing short-lived globs of structured data back and forth, as with an API, I think JSON's a much better fit, as is Protobufs for more tightly joined code.
[1] https://web.archive.org/web/20030228034147/http://www.crockf...
Not really. Plenty of things suck at their original intended purpose and remain in use because they are very good for some other purpose. (Viagra is an example well-known to popular culture, but hardly unique.)
> XML was vastly overused for a long time. That doesn't make those usages correct, as there were alternatives even then.
XML is perhaps not abstractly ideal for many of the purpose it has been used for, but in many cases it was superior to other alternatives for practical reasons, particularly the tooling ecosystem. (JSON is the new XML, and virtually the same thing can be said for JSON in many of its current uses, though it does clean up XMLs two biggest warts, element/attribute distinction and verbosity—though even for human readable formats YAML does the latter even more than JSON while being easier to read, not to mention all the binary options when readability isn't a concern.)
Compare that to JSON or TOML, which are more human-friendly and waste fewer bytes on structure to convey the same information. When used for data storage, two XML files of the same schema describing two completely different objects are likely to share a large amount of content, which is wasteful and gets in the way.
There are many better reasons to hate XML.
And then there is the EU-wide absurdity of WhateverAdES, which invariably leads to onion-like layers of XML in ASN.1 encoded as base64 in XML wrapped in CMS DER encoded message...
That’s easy: S-expressions. Steve Yegge wrote about this in 2005: https://sites.google.com/site/steveyegge2/the-emacs-problem
Would you rather have:
<?xml version="1.0" encoding="utf-8" standalone="no"?>
<!DOCTYPE log SYSTEM "logger.dtd">
<log>
<record>
<date>2005-02-21T18:57:39</date>
<millis>1109041059800</millis>
<sequence>1</sequence>
<logger></logger>
<level>SEVERE</level>
<class>java.util.logging.LogManager$RootLogger</class>
<method>log</method>
<thread>10</thread>
<message>A very very bad thing has happened!</message>
<exception>
<message>java.lang.Exception</message>
<frame>
<class>logtest</class>
<method>main</method>
<line>30</line>
</frame>
</exception>
</record>
</log>
or: (log
'(record
(date "2005-02-21T18:57:39")
(millis 1109041059800)
(sequence 1)
(logger nil)
(level 'SEVERE)
(class "java.util.logging.LogManager$RootLogger")
(method 'log)
(thread 10)
(message "A very very bad thing has happened!")
(exception
(message "java.lang.Exception")
(frame
(class "logtest")
(method 'main)
(line 30)))))Although tbh I'm preferring the XML in your example due to the lack of random quotes - why does (record need a quote, but (log doesn't?
The XML seems a lot more consistent.
Then again: I’d usually only ever choose between formats with support already on the platform/standard library, if it was just some config or small data file. If the data is core to the product then of course it might be reasonable to include a new library or even write a parser. I’m talking java and .net now mainly.
<log>
<record date="2005-02-21T18:57:39" millis="1109041059800" sequence="1" level="SEVERE" class="java.util.logging.LogManager$RootLogger" method="log" thread="10" message="A very very bad thing has happened!">
<exception id="123" message="java.lang.Exception" class="logtest" method="main" line="30" />
</record>
</log>
I also don't object squeezing the exception attributes into record with some prefix that would make the names unique, like that: <record date="2005-02-21T18:57:39" millis="1109041059800" sequence="1" level="SEVERE" class="java.util.logging.LogManager$RootLogger" method="log" thread="10" message="A very very bad thing has happened!" exc-message="java.lang.Exception" exc-class="logtest" exc-method="main" exc-line="30" />
Or, if it's possible to normalize these data, factor out the exception and other attributes and build an index of them so that they can be referenced by an ID: <exception id="123" message="java.lang.Exception" class="logtest" method="main" line="30" />
<record ... exception-id="123" />I'm not being facetious, this is an honest question. Where are the "right" places to use XML?
There are points of disagreement between me and the author, although I wouldn't get too passionate about them.
Super-short version, reading over it again, is that XML is very good at what it does, but it really ought to be seen as a relatively specialized data format. It's really good at certain tasks, best-of-breed for a couple of them, and degrades rapidly as you get away from that. JSON is a fairly cheap & fast general-purpose format that's OK at a lot of things, isn't necessarily great at much, but as you get into more specialized use cases, also tends to degrade. Being a general-purpose format, perhaps arguably it degrades more "slowly", but it does degrade.
Properly understood, IMHO, their use cases don't overlap much if at all, and the combination of them may cover a lot of space, but are still far, far from the only serialization formats you'll ever need.
Why use XML when INI files are way easier to read, and especially edit if needs be?
Why add a ton of additional sugar on top of something as simple as "a=b" for config files?
An example of XML config we used at my workplace was a processing pipeline with various modules and options/parameters encoded for each phase (some optional) of the pipeline. So in a sense it was configuration that resulted in executing code modules, not so much your standard options.
Tim Bray, co-editor of the XML spec, writing in 2006 on the topic:
> Use JSON: Seems easy to me; if you want to serialize a data structure that’s not too text-heavy and all you want is for the receiver to get the same data structure with minimal effort, and you trust the other end to get the i18n right, JSON is hunky-dory.
> Use XML: If you want to provide general-purpose data that the receiver might want to do unforeseen weird and crazy things with, or if you want to be really paranoid and picky about i18n, or if what you’re sending is more like a document than a struct, or if the order of the data matters, or if the data is potentially long-lived (as in, more than seconds) XML is the way to go.
* https://www.tbray.org/ongoing/When/200x/2006/12/21/JSON
He was also editor of the JSON RFCs:
[1] https://en.wikipedia.org/wiki/Darwin_Information_Typing_Arch...
1. SGML (Simple Generalised Markup Language) came first. 2. HTML was a specialisation of SGML, it took off because of the web, and is probably the only reason for SGML to become famous. 3. XML was then invented as a generalisation of HTML, perhaps by people who had never heard of SGML.
And I seem to remember DocBook is an SGML thing, it was invented between steps 2 and 3.
> DocBook is general purpose [XML] schema
[1] http://docs.oasis-open.org/docbook/docbook/v5.1/os/docbook-v...
Also: SGML = standard generalized markup language
I mean, in theory, you could do this in JSON or some other data structure. But you would go insane and be shooting yourself in the head before long.
I'm not sure you could. For example, in another comment, I mentioned DocBook[1]. How would you do the following sample document in JSON?
<?xml version="1.0" encoding="UTF-8"?>
<book xml:id="simple_book" xmlns="http://docbook.org/ns/docbook" version="5.0">
<title>Very simple book</title>
<chapter xml:id="chapter_1">
<title>Chapter 1</title>
<para>Hello world!</para>
<img src="hello.jpg"/>
<para>I hope that your day is proceeding <emphasis>splendidly</emphasis>!</para>
</chapter>
<chapter xml:id="chapter_2">
<title>Chapter 2</title>
<para>Hello again, world!</para>
</chapter>
</book>
Would you make each <chapter> into an object? But you have 2 <para> children in there with an <img> in between. And one <para> has an additional <emphasis> in the content. I can't think of a good JSON schema equivalent to this. {
"id": "simple_book",
"title": "Very simple book",
"chapters": [
{
"id": "chapter_1",
"content": [
{ "type": "title", "value": "Chapter 1" },
{
"type": "para",
"content": [
{ "type": "text", "value": "Hello World!" }
]
},
{ "type": "img", "src": "hello.jpg" },
{
"type": "para",
"content": [
{ "type": "text", "value": "I hope that your day is proceeding " },
{ "type": "emphasis", "value": "splendidly" },
{ "type": "text", "value": "!" }
]
}
]
}
]
}which is probably the only way to properly deal with markup and especially commented sections that can span over paragraph start/ends - neither JSON or XML seems to have a proper answer for such annotations and I wonder if there's any standard format that can that, especially if humans still want to reasonable be able to view or edit iit...
(OOXML and its binary equivalents more or less solve this by completely separating paragraph and character formatting, both separately indexing the spans of text they annotate)
['book', {'id': '...'},
['title', {}, ...],
['chapter', {'id': 0},
['title', {}, 'Chapter 1'],
...
],
['chapter', {'id': 1},
...
],
]
more realistically, you could just represent it with the AST of that XML, i.e {
'type': 'book',
'attrs': {'id': ...},
'children': [
{
'type': 'title',
'children': ['Simple book']
},
{
'type': 'chapter',
...
},
{
'type': 'chapter',
...
},
...
]
}
so you could do that emphasis bit as [
'this text nees more',
{'type': 'emphasis',
'children': ['emotion']},
'!'
]
hellish to write by hand but probably okay for a program to consume (modulo all the XML libs/tooling you can't use). and you could probably even write some kind of schema for it.if i actually had to represent that data, i'd also move some child nodes into attributes, e.g. make all nodes with 'type': 'book' also have a 'title' attribute, like you would if you had an AST datatype
This AFAICT was actually why SVG has a few bizarre choices - such as putting all the drawing commands into attributes. A browser that doesn't understand an embedded SVG document in its HTML would be left with just the text contents.
The include syntax doesn't get enough love. It's crazy that JSON doesn't support it.
Random applications that read XML just aren’t going to implement fully validating parsers, because it is lot of completely unnecessary work. Also in the article mentioned pattern of storing everything in attributes mostly comes from the fact that working with CDATA nodes in XML is major PITA wrt. whitespace handling and coalescing adjacent nodes.
Markup consists of two things: a scalar (a string usually, but can be a binary sequence) and associated structured data: smaller scalars, records with fields, and lists (there's no "etc." here, that's all).
The structured data are either discovered in the scalar by parsing or added to it by marking it up. Parsing applies to binary data and artificial languages (although there are parsers for natural languages as well), marking up to structures that cannot be parsed out, but can be added manually, usually during authoring, but also during after-the-fact indexing.
XML stores both the original scalar and the structure together in a single piece. There's extensive tooling for processing the result.
Practical examples:
1. Parse a C file and do something else with it than compiling. E.g. you want to publish it, index with cross-references, transform maybe: XML shines here (you'll normally want to add XSLT to it).
2. Author text and do something with it. If it's Markdown, apply a minimal parsing and save the resulting AST in XML. Same for reST and any other format out there: just get it into XML as soon as you can and process the XML from that point. Whatever you want to produce (XML, man pages, PDFs), XML toolchain will help you to get there.
3. Mark up existing text. E.g. you have a collection of letters and want to index all references to people. XML would be a very good choice here too. (I'd say that marking up and indexing all existing texts of the humanity would be a very important project. There's already a lot of effort to publish them, and marking up and indexing is what naturally comes next.)
4. I'd venture to say that even binary formats would benefit from conversion to XML and back because of what's possible with XML toolchain (I'm thinking mostly about transformation, but indexing would also be good.) E.g. read a collection of MP3 files, parse them out into what they have (ID3 tags of different versions at the beginning or end, APE tags, other such tags, and MPEG frames), and then do what you want: index by anything, clean up, add extra information that cannot be expressed in tags (classification for classical music or argentine tango, for example) and so on.
PS: Since XML can store structures alongside a scalar, it can also store structures alone: just drop the scalar. It's a very good format for structured data, absolutely not as bad as it's usually painted. Much better than JSON, actually. But you have to prepare it well.
PPS: Scalars and structured data are, of course, the natural parlance of all other programming languages out there, so everything XML does you can do without XML. But it also means that XML is not as foreign as it appears. There is some friction between getting data out of XML and putting it back, but it's about same as with SQL.
That being said, one subtle and important (and often overlooked) difference between XML and, say JSON is that you can stream XML while parsing it on the application level, whereas JSON can not be parsed by the application due to arbitrary ordering of keys. (Of course lower level parsers use streaming anyway, but that's not the point)
In fact you not only can but you should parse XML while streaming it. This is another common abuse: wherever you look you see some high level function that loads an entire parsed XML structure into memory at once. But once you start asking yourself where the file may be coming from you realize that your system may be open to denial of service attacks. E.g. is your system ready to receive a 16GB XML file?
I remember someone describing a trick where instead of sending a JSON array they'd just send a stream of JOSN objects, one per line, so that the receiving end could parse the data in a streaming fashion. But that's not JSON anymore.
- XML schemas give you a ready-to-use format to describe, restrict and document available configuration settings. The unique keys help and a libxml2 gives ready to use validation, even if you may need to 'translate' its error messages before showing them to end users
- XML schemas also support other annotations so you can further generalise your configuration readers by recording the necessary bindings in the XML schema itself, allowing to use it eg to define application user interfaces.
- Almost any text editor can do basic syntax validation preventing most typographic errors, and even better if they can read the schema
- XML schemas are extensible using <import>s, but namespaces still enforce some separation. You can define explicit points where plugins extend your configuration format using <any>
- Human editable - closing tags are noisy but more readable than }],{}] when non-programmers may have to edit these files just to add a few extra textfields to an UI.
- Better datatype support, eg datetimes, by using XML schemas. JSON's type support is too limited
- Support for comments!
- And once you've verified the schema... CSS selectors and DOM APIs to actually process the XML documents.
YAML fixed quite a few things, but still no date times or as far as I know standardised approaches to defining schemas. And I've lost count at how any attempts exist to add schema information or namespacing to JSON...
But for markup... we may be better off to just use markdown inside CDATA blocks
As someone that likes and uses s-expressions, I never thought I would find myself defending XML, but here we are in 2019, no one understands basic parsing theory anymore, and file formats have "evolved" to hot garbage like YAML and TOML.
XML has some great tools in comparison:
http://xmlsoft.org/xmllint.html https://relaxng.org/
> But for markup... we may be better off to just use markdown inside CDATA blocks
SGML can still make a comeback: https://leancrew.com/all-this/2014/09/sgml-nostalgia/
That was the rationalization, in reality I don't know what actual writing use case they were optimized for (Notepad?). XML in a syntax-aware editor with the help of automatic schema validation makes it far easier to write than YAML or TOML.
If your going into complicated things like schemas or other complicated structures, then you probably shouldn't be using YAML or TOML. I would mostly use it for config or other simple things.
At this point if you are not interacting with 20 year old java software, that was created when JSON didn't exist and XML was king, you should be using TOML for simple config, JSON for most things and heavyweight XML, protobuf or csv for the specialized cases. And while we are at it, markdown for simple documentation.
Even gradle decided to use groovy scripts as their config language because it's far more human readable and usable.
You should be using S-expressions.
And once we've just put simple XML file there for configuration and we're past the prototyping phase... "well this actually works good enough, let it be".
Hey, you should checkout http://sgmljs.net (my project).
And of course JSON is the easiest thing to parse in scripting languages, like Python and Ruby -- the entire API is one line, and then you have a native structure you can work with.
It's okay now that the newest revisions support more realistic use cases, but ironically I find it impossible to write as JSON... I write them as YAML which my validator supports natively :)
It's still not as nice as RELAX-NG's compact schema format though IMO :)
https://relaxng.org/compact-tutorial-20030326.html
And the LXML library for Python was a pleasure to use compared to any other similar library I've used.
No we arent't. markdown itself is literally specified as a shortform of HTML [1], and can be translated into canonical angle-bracket syntax using SGML short references (though not completely eg. markdown reference links require unlimited forward lookup). This gives a canonical representation of markdown in SGML/XML even if you don't use SGML.
(4 lines, 75 characters)
> Here's the right way:
(10 lines, 133 characters)
I have a suspicion as to what went wrong.
The author's point is that XML should not be a data format.
The point being, I think the author is arguing you should use the right tool for the job, and XML not being designed for arbitrary data structures makes it not the right tool. Just recently people have shown you can build a raytracing engine in SQL, but if someome was arguing we shall call it SQLCycles and ship it in Blender, I'd definitely have a few objections!
For humans, this is documents with certain meanings attached. For computers, it’s documents with certain meanings attached.
It’s all data. XML is a data markup language. It’s just that humans call it “semantics” in a “document.”
It easy to query with xpath
It parses the same regardless of the order. With the right way, is <item><key>Name</key><value>John</value></item> the same as <item><value>John</value><key>Name</key></item> ?
<key>CFBundleDisplayName</key>
<string>TextEdit</string>
<key>NSHumanReadableCopyright</key>
<string>Copyright 2019</string>
This may be perfectly parsable by a SAX parser storing some state, but its totally not processable by xslt.
I had the impression mobile was some new try at frontend tech, but somehow iOS and Android threw a whole bunch of outdated stuff at me.
I mean, before the iPhone I designed a XML based ETL config system and tried to avoid all the common XML errors, then I start doing a mobile app 10 years later and it's like all that knowledge was forgotten...
It's a bit inconvenient but perfectly processable with XPath:
/dict/key[.='CFBundleDisplayName']/following-sibling::string[1]And yes, plists suck and make your XPath selectors ugly, although you could write a function to abstract them out.
I like this way
<root>
<item key="name">John</item>
<item key="city">London</item>
</root>
So I can use this xpath to get the person's name: //root/item[@key="name"]/text()
Not sure what would be the xpath to get the name if the XML was <root>
<item>
<key>Name</key>
<value>John</value>
</item>
<item>
<key>City</key>
<value>London</value>
</item>
</root>
This is a better example: <employees>
<employee id="1">
<field name="name">John</field>
<field name="city">London</field>
</employee>
<employee id="2">
<field name="name">Jack</field>
<field name="city">Boston</field>
</employee>
<employees>//root/item/key[text()="Name"]/../value/text()
//root/item[key[text()="Name"]]/value/text() //root/item[key="name"]/value <employees>
<employee id="1">
<name>John</name>
<city>London</city>
</employee>
<employee id="1">
<name>Jack</name>
<city>Boston</city>
</employee>
I know what you’re trying to do there - you’re trying to “future-proof” your schema by allowing introduction of arbitrary new elements. Which means that there’s no standard way to guard against somebody omitting a required field (like “name”) or adding a new field like “creditCardNumber” - other than to document your acceptable key values in a non-standard format and add defensive code that a validating parser would have given you. You’re better off taking as much advantage of the format as you can.While XML can seem cumbersome (compared to JSON say) it is a very good 'data transport' tool when used correctly with a sensible schema (XSD).
For example, we use XML as a 'vendor neutral' data format to export/import CAD geometry and associated data for town utilities such as buildings, pipes, roads etc. All this data has to be validated against the schema to ensure its correctness. Using a schema like this enables the city council to import this XML into the GIS system to be used for asset management, financial planning etc.
A good schema can be key to sharing XML effectively between departments/applications and being a markup language this data can also be viewed independently using XLST.
The correct way to express a dictionary in XML is something like this:
<root> <item> <key>Name</key> <value>John</value> </item> <item> <key>City</key> <value>London</value> </item> </root>
In the past I used to create scripts that exported xml from relational data but didn't really understand the right way to build and structure them.
If you told me that the transmission and parsing rate is too slow for their application, that's a real dig at it.
Almost.
The idea, for those not familiar, is that once a work of art is published (a novel, a poem, a song, a painting), it speaks for itself, and authorial intent no longer matters.
That is, meaning and purpose are in the eye of the beholder/consumer. And there is no right or wrong way to "interpret" art. If someone finds meaning that the author did not intend, it is just as valid as a deeply hidden but intentional allegory they intentionally placed in when they were writing.
The relevance to software is it applies to APIs, specifications, standards and formats.
There is no such thing as users using your software or specification "wrong" - if they insist on doing so, the meaning has evolved. Evolve with it or die.
XML wasn't an original invention; it is specified as a proper SGML subset. From the XML spec:
> The Extensible Markup Language (XML) is a subset of SGML that is completely described in this document.
Now I totally agree that SGML and XML aren't for service payloads and config files. The sole purpose of markup languages is representing structured text. And arguably, SGML fills this role much more adequately than XML today as it can represent (via the SHORTREF mechanism) custom Wiki syntaxes such as markdown and others, and in contrast to XML, can deal with the largest corpus of markup out there eg. can parse HTML with all its minimization features such a omitted tags, enumerated and unquoted attributes, etc. See [1] for a practical introduction (disclaimer: link to a tutorial I held last month at ACM DocEng).
<!ELEMENT e - -(f,g,h) -- no tag omission -->
<!ELEMENT f O - (#PCDATA) -- start-tag omission -->
<!ELEMENT g - O (#PCDATA) -- end-tag omission -->
<!ELEMENT h O O (#PCDATA) -- both start- and end-
tag omission allowed-->
What's painful about end-element tag omission?The lack of </> in XML is a crime against humanity.
https://www.egosoft.com:8444/confluence/display/XRWIKI/Missi...
This is the language used in the game X4 Foundations (and others in the series). An example of its use (mine, i claim no grace in it):
https://github.com/h2odragon/dragoncommands/blob/master/aisc...
... XML is not a great format for an extension language, I have to say.
It basically names the tag what you name the node. :S
- You link/unlink nodes (I called them entities! Xo) by right-click-dragging between them.
- You copy stuff by right-click-dragging to an empty space.
- You delete by grabbing something by left-click-holding and pressing the delete key.
- Oh, and nodes are completely tree structure expandable, just drag-drop attributes on nodes and nodes inside nodes.
The editor uses lightweight rendering so you can have a ton of elements with good performance.
(I know, not super intuitive; but very handy once you know about these.)
Magento 2 (acquired by Adobe for $1.68bn) uses XML to render its layouts. Here's some fun XML for the checkout page:
https://github.com/magento/magento2/blob/2.3-develop/app/cod...
It had to survive power outs too.
<RequiredPackages Count="5">
<Item1>
<PackageName Value="LazUtils"/>
</Item1>
<Item2>
<PackageName Value="treelistviewpackage"/>
</Item2>
<Item3>
<PackageName Value="internettools"/>
</Item3>
<Item4>
<PackageName Value="LCLBase"/>
<MinVersion Major="1" Release="1" Valid="True"/>
</Item4>
<Item5>
<PackageName Value="LCL"/>
</Item5>
</RequiredPackages>https://leancrew.com/all-this/2014/09/sgml-nostalgia/
SGML really needs a revival.
Software v1.0: Config is a binary blob and everybody curses the nasty editor provided by vendors.
Software v2.0: Config is in XML. Thank god!
Oh wait.
<Binaryblob><Byte value="65"/><Byte value="99"/> ....
<root>
<Name>John</Name>
<City>London</City>
</root>
Removes one level of indirection, XML already has keys.It's also not a good example of XML, because XML schemas should have a fixed list of tag names.
<item name="name" value="John" />
<item name="city" value="London" />
where the key names are used as attributes, so it wouldn't work with arbitrary key names either, right? <dict>
<key>CFBundlePackageType</key>
<string>FMWK</string>
<key>CFBundleShortVersionString</key>
<string>4.8</string>
...But you answer your question, yes this format can be expressed with XML-Schema:
<xsd:element name="dict">
<xsd:complexType>
<xsd:sequence minOccurs="0" maxOccurs="unbounded">
<xsd:element name="key" type="xsd:string" />
<xsd:choice>
<xsd:element name="string" type="xsd:string" />
<xsd:element name="integer" type="xsd:integer" />
<!-- etc. -->
</xsd:choice>
</xsd:sequence>
</xsd:complexType>
</xsd:element> grammar {
start = Dict
Dict = element dict { DictItem* }
DictItem =
element key { text },
(element string { text } | element number { xsd:integer })
}Now, if we're just encoding a dictionary that's an already an encoding of an object, then yeah, let's just encode the object directly like you are above.
Which is why I still use the simple, effective, INI format for configuration files for applications I write.
XML for config files is madness personified.
"XML is like violence – if it doesn’t solve your problems, you are not using enough of it."
https://libvirt.org/formatdomain.html
It does occasionally put information outside of the tags, but because there's no logic to when, it's nearly worse.
There are many, many more inventions that are used for different purposes they were meant for.
The Internet was created so that US can withstand nuclear attack and it was never meant to be primarily used to spread advertisements.
Get over it.