A Brief Defense of XML
borretti.me
borretti.me
You might be thinking, "Duh, just escape it". Nope, ASCII control codes are forbidden you can't escape them, the escaped form is still forbidden. You can escape say, a backslash, or a quote mark, or an angle bracket, but you can't escape say a vertical tab.
So if you want to move arbitrary ASCII text (or Unicode text, since ASCII is a subset of Unicode) with XML you need to Base64 encode it, or you won't succeed.
15-20 years ago I wrote code which needed to square this particular circle where clients insisted that they want XML text and yet they also demand to handle the forbidden values, my code would just convert all impermissible values to Unicode's U+FFFD the Replacement Character.
> Char ::= [#x1-#xD7FF] | [#xE000-#xFFFD] | [#x10000-#x10FFFF] /* any Unicode character, excluding the surrogate blocks, FFFE, and FFFF. */
I’m puzzled at the total non-mention of U+0000 in the document, which is a valid Unicode and ASCII character but is not supported by this definition of Char. So it seems like it’s actually (and using better terminology) “any Unicode scalar value, except U+0000, U+FFFE and U+FFFF”.
The same restriction you noted on entity encoding still holds: https://www.w3.org/TR/xml11/#NT-CharRef
Control codes aren't text. Heck, HTML doesn't even support tabs -- whitespace is just generic whitespace, interpreted as a regular space.
If you want to insert non-typographic elements that affect the layout/presentation, that's what <elements> are for.
While if you need to encode binary data that can represent ASCII text, then yes -- Base64 is an ideal solution. That guarantees it won't be garbled by other XML processors.
The W3 basically says all this here [1]. ("Use of control codes in HTML and XHTML is never appropriate, since these markup languages are for representing text, not data.")
Sure it does. As does XML. You just have to choose deliberately to use it, via such means as <pre> or the CSS `white-space: pre`, or, in XML, the xml:space attribute.
Incidentally, HTML and the web is a place where the superiority of tab-based indentation to space-based is particularly apparent, because you’re far more likely to encounter very narrow viewports, where reducing the width of tabs is useful, while still keeping a reasonable indentation level on wide enough screens. I do something similar to this, for example:
:root {
tab-size: 4;
@media (max-width: 50rem) {
tab-size: 2;
}
}You could also use tags, for example <br> tag in HTML.
After all it is Extensible Markup Language and the functionality for adding custom tags is the basic feature.
And as suggested above, base64 encoding and application level decoding is also an accepted practice.
All these techniques point to handling the rendering at the application level.
I feel seen. I remember suggesting this during a design discussion where the issue was how to embed an arbitrary document in another, where the embedded document could even be ill-formed. Everybody told me it was a gross and dirty hack and ignored the suggestion.
So, if one is building an application and needs to render a proprietary data object, this is a good way (not the only one) to handle it using a client library.
Is XML meant to represent only text and not data?
It's easy to understand why -- pre-XML, data exchange formats were generally binary or otherwise non-readable. But once storage got generous enough that each byte stopped mattering, people looked for a textual data format that was easy to debug, and XML was available.
Still, it does make we wish a happy middle ground could have been found. An open standard something essentially like a compact binary format (similar to protobufs) but with a schema with field names and defined structure required at the top of every file. Which would require a dedicated app to view it, essentially similar to JSON once parsed. (While not transmitting the schema during communications but hard-coding it into the sender and receiver could serve as a poor man's obfuscation.)
If you want to see how to do these XML-to-JSON and JSON-to-XML transforms I have written a little learning repo with a CLI: https://github.com/aleph2c/leaning_xslt
Here is Michael Kay's white paper on Transforming JSON using XSLT 3.0: https://www.saxonica.com/papers/xmlprague-2016mhk.pdf
Once your data is in a JSON format, you could implement your compact-binary-format idea around it.
If we are authoring content that we expect to last a few decades, markup is the proven engineering choice. It used to be sigil markup $1, $x, @@p etc., and other proprietary markup. SGML came and introduced DTDs, entities, parsing etc.,
DocBook format is pretty good for technical manuals. They pretty much covered 98% of the use cases for technical books. It would be foolish to invent a different format (using say JSON representation or some other proprietary format) for those use cases.
SGML grew too complex and then came XML in the late 90s and took off when SOAP started gaining traction for web services in the early '00s.
Unfortunately, XML almost went the SGML way on complexity with the addition of namespaces, multiple schema languages etc., But, we can never fault XML for its expressivity. There is also a good tooling available due to the work of many standardization workgroups and browser/vendor support.
If you choose to archive content using XML, you can be guaranteed that you can parse it 20-30 years later since XML is a simple format (if you use a simple subset of features).
With this said, Michael Kay offers a nice series of patterns in chapter 17 of his book "XSLT 2.0 and XPath 2.0 Programmer's Reference". He shows how to use a "Fill-in-the-blanks" stylesheet that someone who only understands HTML can use, then he shows how to make a "Navigational Stylesheet", a "Rule-Based Stylesheet" and finally a "Computational Stylesheet". He shows that you don't necessarily need to use XSLT's functional programming powers, how you can get a lot with a little of it, and how to level up. So it seems like the language was designed to be written by people who don't understand it; it's declarative.
Like you said - standard XML isn't terrible. Adding on an XSD isn't terrible, because now you can enforce structure and datatypes on files provided by outside parties. Creating an XSLT is much more of a mental challenge, and probably should be left to tools to define.
Anything beyond those technologies is someone polishing up their resume.
bayerName="Beta Pictoris"
spectralType="A6V"
mass="1.75"
luminosity="8.7">
<position
rightAscension="05h 47m 17.1s"
declination="−51°03′59""
distance="63.4ly"/>
<age min="20My" max="24My"/>
</star>The tag-content ratio of XML is only as bad as you make it.
Admittedly, XML isn't the most pleasant syntax to interact with by hand.
The main idea, however, was wonderful. An information exchange format (xml itself), a way to define schemas (xsd), a transformation language (xslt), and stylesheets (xsl?). I kind of miss it.
XSLT and XPath each went on to great solo careers, XSL-FO kind of fizzled when XML+XSL failed to displace HTML+CSS.
There's also XQuery for transformations and XSD/RelaxNG for schemas, the whole combination of XML database, WYSIWYG (via XSLT) XML editor and output transformations is very powerful, every journal we published and every site we ran at the British Medical Journal was generated that way.
A lot of really stupid JavaScript could be done fairly easily with XSLT natively inside a browser engine. An XML document can point to its own XSL(s) so get rendered to the appropriate delivery format or just used directly as data.
No, I'm pretty sure that if they have only used XML syntax as a data rep and not XML-associated tooling like XSD, they are quite aware that when thet say they hate XML they are talking about syntax.
OTOH. I am surprised at the number of people that seem to thing “has a schema language that can be used to specify structure beyond just the basic raw language” distinguishes XML from other data representation languages.
Or, more generally, the way that “if you don’t like XML, maybe your problem is you aren’t using enough of it” has become an edgy contrarian position rather than steretypical establishment enterprise consultant thing it was 15 or so years ago.
JSON has .avro, I thought.
It's all data. Validation logic is sold separately.
JSON, like XML, has several schema languages, which have widely-available libraries for generating host language typings, validation, etc.
> Typing in XML is far richer than in JSON. There’s even a date type!
JSON Type Definition has timestamp, JSON Schema has date, date-time, time, and duration.
{
"$schema": "#/$theSchema",
"$theSchema": {
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "MySchema",
"type": "object",
"properties": {
"productId": {
"description": "The unique identifier for a product",
"type": "integer"
},
"productName": {
"description": "Name of the product",
"type": "string"
},
"price": {
"description": "The price of the product",
"type": "number",
"exclusiveMinimum": 0
}
}
},
"productId": 1,
"productName": "An ice sculpture",
"price": 12.50
}I've seen JSON schemas describe things like date types and plenty more. If you are including schema metadata alongside XML, then for a fair fight you can include JSON schemas, too.
A particularly unusual take is to ignore the difference and allow either attributes or child elements as different ways to do the same thing. This is what XAML does; e.g. this:
<Button Content="Foo" />
is identical to: <Button>
<Button.Content>Foo</Button.Content>
</Button>If data should be presented to the user by a user agent that doesn't understand extensions to the markup, it should be using text nodes.
If it is a modification to a subset of a block of text, such as emphasizing a word, it should be a new element.
If it is a block of semantic metadata, you would structure that as a hierarchy of elements with attributes.
Otherwise, it is usually an attribute on an existing element.
SVG is a bit of a zany format in XML because so much of the content is points or encoded data, which get stuffed in some truly epic attributes.
XML wasn't really made for pure data formats, even though shoehorning that in was the majority of the market's effort.
Using plain tags is so obvious that just about anyone can get it, and in my experience there's no real advantage to using attributes.
At least that's my experience at work, where we use XML as our primary integration format for others that want to deliver data to our system. When going through with a new customer, the few attributes we have almost always cause some kind of issue or misunderstanding.
In the text markup role tags are like functions and the attributes are parameters to the functions. but in the data encoding role the attributes are another redundant key value store and you end up with a confusing mix where some of your kv's are in the tag and some are in the attributes and it is a bit of a mess.
There is a star tag(presumably a list of star tags) each of which contains a single position tag. The star tag is is full of attributes, and the position tag is full of attributes, now you have a complicated(not that complicated, hyperbola is for arguments sake) access method where you need to do different things based on what data you want out of the structure. Why was the position tag not an attribute? What happens when you want to encode the rightAscension, and declination, they will get their own tag, in fact they probably just should have tagged everything in the first place and avoided attributes.
So then:
<star
bayerName="Beta Pictoris"
spectralType="A6V"
mass="1.75"
luminosity="8.7"
rightAscension="05h 47m 17.1s"
declination="−51°03′59″"
distance="63.4ly"
minAge="20My"
maxAge="24My"/>
We are, of course, playing into the article’s second footnote here.(Also: I’m curious why you turned the original document’s ″ (U+2033 DOUBLE PRIME) into " (", U+0022 QUOTATION MARK).)
I don't ever see double prime at my job, so it was probably muscle memory.
I wonder when XSLT will be re-invented and applied to JSON and YAML. (XSLT was a bit of a miracle: a pure functional language adopted by a significant chunk of non-CS community in early 2000s. The syntax forced by XML sucked though. YAML can do better!)
(Could use S-expressions as well. The approach is 60+ years old now...)
Are you talking specifically about XSD and JSONSchema? If so I don't think this is really true. Pretty sure the original authors of JSONSchema didn't pay much attention to XSD.
So for a lot of XSD's I'd say JSON Schema is effectively very close.
Sadly, tooling for the latest JSON Schema (which is required for choice elements IIRC) is lacking, Swagger UI for example does not work well with it for example.
[1]: https://json-schema.org/draft/2020-12/json-schema-core.html
Can it do the reverse?
I don't see any big issues with doing it in reverse for a subset of JSON Schema.
What I haven't looked at is how large that subset would be. That is, if JSON Schema contains things that aren't possible to represent in XSD.
For example, JSON Schema allows for some extensive conditional sub-schema logic. I used the "oneOf" to represent a XSD choice element, however JSON Schema also allows for such things as "if..then..else", which as a self-taught non-expert on XSD's I'm not sure how would map to XSD on a Saturday morning.
Of course, there are things in XSD which can't be represented in JSON Schema either, I'm sure it's easy to come up with something that doesn't map. But from the typical XSD's I've seen and the ones we had in our company, the mapping could be done with just a single special-case.
Given the lack of tooling around JSON Schema, my workflow now is to create/update the XSD and regenerate the JSON Schema.
I really don’t think YAML and its own terse but ambiguous rules are better suited. The problem I’ve found in XSLT during this effort, and in YAML in the past, is there’s no debugging story that isn’t full of pain and mysterious grasping in the dark.
I know actually functional languages like lisps do this better. I strongly suspect more declarative logic languages do better still. I also know from experience that the Excel interfaces most of our users interface are better than the XML experience, and I strongly suspect that’s closer to the pure/declarative + non-programming-accessible ideal than a markup language suited to other purposes.
That's jq and its gojq friend (for its `--yaml-input`), right? If not a standard, damn near one as best I can tell. JMESPath (used by awscli and ansible) is a joke compared to the expressiveness of jq. There's also JSONPath in kubectl but it is similarly nowhere near as awesome as jq and not even complete in its attempt to be "XPath for JSON"
But I don't think anyone is arguing that it's completely useless.
It just hasn't earned it's current privileged position in the computing world. It got there on the coattail of the obscene wave of hype around the turn of the century.
So it has a way to go as it slowly falls back to its rightful place.
Before computers, you could still reasonably use the term document to refer to things as different as an essay, poem, book manuscript, CV, patent, birth certificate, voter-registration form, and a court transcript.
It would be nice to cleave them apart, but that'll take a lot of work since our human languages have a dearth of good terms for discriminating between highly-structured and mostly-unstructured documents.
I struggled through writing something roughly about this last year: https://t-ravis.com/post/doc/what_color_is_your_markup/
In those terms, the markup depicted in the linked piece's "document" is basically all "structural" markup--and I feel like this is the least interesting part of an unstructured document. It's the same logic that translates a paper address-change form into structured data describing what's in the fields--but it doesn't yield any of the leverage that we get from the meaningful associations in the form-as-markup.
We can't do much more with the structural markup of a free-form document than present it and perhaps translate it to other formats with similar idioms. Structure in this kind of document is somewhat capricious (if five different writers wrote exactly the same 20 pages of text, they might all still use different section/paragraph/sentence boundaries).
The bigger leverage in free-form documents comes from ontological markup that annotates what's being written about. This is what's going to enable you to bolt on an interesting extension that your readers can use to jump between every section on your site that discusses the same paper or author. Or enable you to automatically inject birth/death/release dates for people, films, and albums you refer to.
(I don't mean to suggest the post precludes these--but the verbosity comparison feels much less fair without enough inline annotation to provide a similar level of utility as the star record.)
inst:
Type: AWS::EC2::Instance
Properties:
UserData: |
#!/bin/sh
echo "hello, world!"
cat >/root/.ssh/authorized_keys <<FOO
ssh-rsa cafebabedeadbeef
FOO
versus {
"inst": {
"Type": "AWS::EC2::Instance",
"Properties": {
"UserData": "#!/bin/sh\necho \"hello, world!\"\ncat >/root/.ssh/authorized_keys <<FOO\nssh-rsa cafebabedeadbeef\nFOO\n"
}
}
}Compared to "string" and "character," not so much and not to so much misfortune.
> XML is Lisp’s bastard nephew, with uglier syntax and no semantics. Yet it is poised to enable the creation of a web of data that dwarfs anything since the Library at Alexandria.
From A Brief Defense of XML:
> ..Despite its roots in SGML (the Common Lisp of markup languages), the creators advertised it as a general format to exchange any digital information.
From The Nature of Lisp: https://www.defmacro.org/ramblings/lisp.html
> All this effectively means that we can use XML for generic storage of source code. We'd be able to create a whole class of programming languages that use uniform syntax, as well as write transformers that convert existing source code to XML. If we were to actually adopt this idea, compilers for different languages wouldn't need to implement parsers for their specific grammars - they'd simply use an XML parser to turn XML directly into an abstract syntax tree.
> ..Everything we've learned about Lisp so far can be summarized by a single statement: Lisp is executable XML with a friendlier syntax.
They solve concrete problems and if one wants to start fresh with something like JSON they need to be reinvented.
Once all the functionality is replicated its quite likely that the cognitive difficulty of mastering it all is - more or less - the same.
Lets also not forget that tech in all its forms has been quite democratized in the last decade. Corporate and academic environments do not always care about removing needless hoops and may even have incentives to keep things arcane and less accessible
Trying to map it cleanly to a strongly typed world was the problem.
The idea of using XML as HTML was... an idea, but rather than trying to rationalize it as a super complicated dtd, they should have instead opened it up and defined mechanisms for creating custom tags, sort of like how React works now.
XML is great for applications where you might want to discard some or all of the data easily; shame the way it was employed required you to pretty explicitly handle every possible permutation of the data.
It was still kind of verbose, but effective.
Eventually switched to Markdown+Pandoc+more custom scripts. Formatting Objects and Apache FOP seemed to be on the back burner anyway.
I like XML for markup in theory, and the formality is nice, but, damn, is Markdown nicer in practice for documents.
XML got a bad rap because SOAP & enterprise built nightmarish systems. It's possible to do very bad things with any kind of data interchange system, which, contrary to this post, I would say XML is & does fine at.
Don't really have anything bad to say about it.
"Web framework"? Hahaha!
UI (HTML/etc) was produced by XSLT at request time; such JS as we used was our own; Apache/FastCGI; uhhhh.
yaml, json, xml all seem to be disliked.
ini works, but is pretty limited. It's also likely disliked, but just (at this point) obscure.
The side of the fence you're on is directly proportional to how old you are. Neckbeards who want you to get off their lawns love XML (related: get off my lawn). And the "kids" will give a bajillion reasons why we are "technically" incorrect...
If you split hairs here... You've never lived in a world without XML (truly terrifying TBH).
Neckbeards who want you to get off their lawns hate XML and will wax lyrical about how S-expressions endow data with type semantics, whereas XML just gives tree structure to pieces of text.
Yaml specification makes Xml specification look like JSON.