Not the comp.text.sgml FAQ (2002)
flightlab.com
flightlab.com
Q. I'm designing my first DTD. Should I use elements or
attributes to store data?
A. Of course. What else would you use?
I giggled.Q. What's so great about ISO standardization?
A. It is often said that one of the advantages of SGML over some other, proprietary, generic markup scheme is that "nobody owns the standard". While this is not strictly true, the ISO's pricing policy certainly has helped to keep the number of people who do own a copy of the Standard at an absolute minimum.
[ Ed. note: I'm not exactly sure why this is seen as an advantage,
it's just something people say. ]Both (elements and attributes) are unsuitable to store arbitrary data in a straight forward manner. If you absolutely must, you should know about normalization, XML-whitespace handling and CDATA.
SGML and XML are for text, optionally marked up with tags/elements. Attributes are for data about element presentation, and not intended to be displayed directly. It's as simple as that.
Usage of XML in business data and non-text document formats OTOH is an accident IMHO, but is still the most robust format for data exchange and archival in long-term commitment scenarios we've come up with so far, and I don't see that changing anytime soon because there aren't that many open standards being developed anymore.
Does SVG offer any advantages over, say, EMF/WMF? I know a strong advantage of WMF is that the file format translates 1:1 into GDI calls, which makes it very fast - but I don't see EMF files rendered with anti-aliasing or complex gradients. What about PDF or PostScript?
Good SVG can even be human-readable and -editable. I've actually fixed simple broken SVGs with vim and a pocket calculator.
PostScript is great but as a binary format not very modern, where text-based formats seem preferred. I had to look up WMF and dumping calls to Microsoft API as an 'open' standards does not sound too exciting, either.
Its disadvantage is that it is a full on programming language, not declarative, so it isn't friendly to editing tools.
Sadly, programming is coming to SVG as well with more and more renderers supporting JavaScript. Which might be great for some use cases, but again even further disperses the field of possible generator/feature/renderer combinations that might (will) fail.
S... strong advantage?
Somewhat of a tangent, but I like the name of the alegraic structure for binary trees: magma [0].
So is ASN.1, but nobody in their right mind would use ASN.1 as a markup language. The sole fact that it is possible to cram the structure into the format doesn't mean it's a good idea.
Here I'm thinking of the correspondence between trees and strings of nested parentheses. So, strings like
(a (b c) d) (e f)
carry a tree structure, and any (non-associative) multiplication is just a fold operation over a tree. In this case it's like we tag each set of prentheses with a name/operation, so naively XML would seem to naturally represent this kind of thing.
I have little actually experience using XML directly though, so am genuinely curious as to what so terrible or "crammy" about my ideas here.
There is some mismatch between what XML was designed for and the problems XML is good at solving. It was most certainly designed to be easy for humans to read and write manually. In practice, it is a great interchange format, by which I mean the specific idea of different parties writing XML for exchange between each other, because it can be validated mechanically. There is a large and powerful ecosystem of software that has sprung up around it which simply isn't there for s-exps or JSON.
Elsewhere on here, marcoperaza points out that it is kind of against hacker ethic. That's true. But enterprise software often involves multiple separate organizations having to agree on what a document can contain, and XML is great for that, and that use case tends to be more valuable in industry than whether it is the tersest, most flexible or readable format.
for example, i often use xml documents with no text - for example
<contraindications> <pregnancy/> </contraindications>
and really the rule of thumb about attributes vs elements is more about if you are going to need chikdren of item - easier to extend elements with child nodes..
Almost all characters are permitted in names, except those which either are or reasonably could be used as delimiters. The intention is to be inclusive rather than exclusive, so that writing systems not yet encoded in Unicode can be used in XML names
I'd say anything which can be nested.
For example, XML makes sense for a layout engine:
Arbitrary amounts of buttons can fit inside a layout? Make it a child.
Text in a button? If only one and can't have children tags? Make it an attribute. Else -> Make a child.
The first one claims (without giving a source) that James Clark once said or wrote:
“Any damn fool could produce a better data format than XML” – James Clark 2007-04-06
Bizzare.
There are plenty of good arguments for the XML way of doing things. For example, having a rigorously defined way (XSLT) to specify transformations of schema-conforming XML is more robust than ad-hoc code that wrangles schemaless JSON.
But it does go against the hacker ethos and stands in the way of rapid development. And wherever it is used, complexity and verbosity seem to often follow. Look at SOAP, for example.
But the xml ecosystem is horrible. Sensible ideas, horrific execution; like namespaces and schema. Probably the single worst problem was using xml syntax itself: it's like, a programming language that uses JSON for its syntax.
But also, there's guilt-by-association, people hate the enterprise culture that uses xml - similar happened to java.
Though xpath is not so bad, and many people seem to quite like it.
Finally... json is a better match for data, basically by being c-like. However, an ecosystem tumour is also growing, around JSON. Some even use json syntax itself...
I wonder, if perhaps, a root issue is that the world is complex, and youthful simplicity is corrupted as it adapts to cope with the real world... There is hope, however; tools like `jq` never existed for xml.
s/serialization format/object notation/This also got me reading up on various structured-data formats: XML, YAML, JSON, TOML, HCL, etc. I'd really like some big table comparing various features but can't seem to find anything of the sort.
I found a link [0] that has comparisons between JSON, TOML and YAML representations for various types of data. It's neat to see how each becomes more or less verbose depending on the kind of data getting encoded.
[0]: https://gohugohq.com/howto/toml-json-yaml-comparison/
edit: I found a table:
https://en.wikipedia.org/wiki/Comparison_of_data_serializati...
XML is fine if all you want is a human-readable format to define tree data structures such as documents to be used in applications where only strings are used and someone within the use case needs to have the semantics of each node and each attribute spelled out quite clearly and unequivocally in the document structure itself.
For any other case, XML is horrible.
Now, consider that XML is used quite extensively in any other case beyond tree-based DOM data structures that it was designed for.