XML/SGML are a very effective way of representing tree data (again, non-1D data) in 1D strings. And they're wonderfully extensible, while still keeping a well defined schema. Is there an alternative language that is better? How is it better?
XML/SGML are a very effective way of representing tree data (again, non-1D data) in 1D strings. And they're wonderfully extensible, while still keeping a well defined schema. Is there an alternative language that is better? How is it better?
If you decided to go fully XML on "<path d="M 10 10 H 90 V 90 H 10 L 10 10"/>" it would probably look like this:
<path>
<point>
<x>10</x>
<y>10</y>
</point>
<point>
<x>90</x>
<y>10</y>
</point>
<point>
<x>90</x>
<y>90</y>
</point>
<point>
<x>10</x>
<y>90</y>
</point>
<point>
<x>10</x>
<y>10</y>
</point>
</path>
Some would say the fact the designers of XML went for "M 10 10 H 90 V 90 H 10 L 10 10" instead shows they thought XML is too verbose.simias likely agrees that XML is too verbose - and wonders why they used XML at all
And it’s not like XML has no good ideas in it either. You may disagree about the execution, but the idea of seamlessly mixing structured data defined by independent standards in a single document is pretty damn powerful.
Also, your example is rather strawman-ish: there is no point in using separate elements for individual coordinates, and anyone designing the format would know this. A more realistic example would have <path> contain an ordered sequence of <move-to x="..." y="..." /> and <line-to x="..." y="..." /> elements. (In fact, SVG already contains <line /> and <circle /> which use roughly this structure.)
The thing is, just having an XML parser isn't enough to parse SVG. You also have to parse a DSL. So I think the argument is "why not just use only a DSL, that does a good job of describing the data model?"
I imagine they thought that the path DSL is simple enough to parse (and I seem to vaguely recall PostScript has something similar), while the overhead of representing path nodes as XML elements would be too high.
I can't help but think s-expressions would have been a better choice.
You say that - but here's some real world XML from the widely used 'GPX' file format[1]
<trkpt lat="47.644548" lon="-122.326897">
<ele>4.46</ele>
<time>2009-10-17T18:37:26Z</time>
</trkpt>
The truth is I could have put the X coordinate as an attribute and the Y coordinate as a child element and it would still have been a fair representation of real-world XML documents.And GPX is one of the better XML formats! You want to see nightmare XML? Go look at SAML.
[1] https://en.wikipedia.org/wiki/GPS_Exchange_Format#Sample_GPX...
JSON and YAML do not make such a distinction.
(I do concur with others who have noted here that the "elements" versus "attributes" distinction makes XML a poor choice for serialization, but XML as a serialization format isn't really the issue here.)
So using XML to represent data structures instead of marking up text can be pretty awkward, inefficient, and nuanced.
XML Attributes are second class citizens compared to TEXT nodes which can contain CDATA, because attributes undergo "Attribute-Value Normalization" -- having their line breaks, entity references, and white space normalized. Newlines are normalized, leading and trailing white space removed, repeating white space replaced with a single space.
SVG path attributes (as well as simple values like numbers, booleans, enums, etc) are impervious to Attribute Value Normalization corruption, because they don't depend on white space being perfectly preserved (by design, of course), so they are fine to put in attributes.
But if you really care about preserving the exact value of a string, like a password or arbitrary string, you should use <!CDATA[[ ]]> in a text node, not an attribute!
I Wanna Be <![CDATA[ https://donhopkins.medium.com/twenty-twenty-twenty-four-esca... ]]>
https://www.w3.org/TR/xml/#AVNormalize
3.3.3 Attribute-Value Normalization
Before the value of an attribute is passed to the application or checked for validity, the XML processor must normalize the attribute value by applying the algorithm below, or by using some other method such that the value passed to the application is the same as that produced by the algorithm.
All line breaks must have been normalized on input to #xA as described in 2.11 End-of-Line Handling, so the rest of this algorithm operates on text normalized in this way.
Begin with a normalized value consisting of the empty string.
For each character, entity reference, or character reference in the unnormalized attribute value, beginning with the first and continuing to the last, do the following:
For a character reference, append the referenced character to the normalized value.
For an entity reference, recursively apply step 3 of this algorithm to the replacement text of the entity.
For a white space character (#x20, #xD, #xA, #x9), append a space character (#x20) to the normalized value.
For another character, append the character to the normalized value.
If the attribute type is not CDATA, then the XML processor must further process the normalized attribute value by discarding any leading and trailing space (#x20) characters, and by replacing sequences of space (#x20) characters by a single space (#x20) character.
Note that if the unnormalized attribute value contains a character reference to a white space character other than space (#x20), the normalized value contains the referenced character itself (#xD, #xA or #x9). This contrasts with the case where the unnormalized value contains a white space character (not a reference), which is replaced with a space character (#x20) in the normalized value and also contrasts with the case where the unnormalized value contains an entity reference whose replacement text contains a white space character; being recursively processed, the white space character is replaced with a space character (#x20) in the normalized value.
All attributes for which no declaration has been read should be treated by a non-validating processor as if declared CDATA.
It is an error if an attribute value contains a reference to an entity for which no declaration has been read.
You should read some of James Clark's criticisms of the official XML Schema (XSD) standard, which motivated him to develop TREX (Tree Regular Expressions for Xml), which he combined with Makoto Murata's RELAX (REgular LAnguage description for XML) to create Relax/NG.
https://en.wikipedia.org/wiki/James_Clark_(programmer)
https://en.wikipedia.org/wiki/Makoto_Murata#RELAX_and_RELAX_...
>Some people, including Murata and James Clark, had critical attitudes toward XML Schema. XML Schema is a modern XML schema language designed by W3C XML Schema Working Group. W3C intended XML Schema to supersede traditional DTD (Document Type Definition). XML Schema supports so many features that its specification is large and complex. Murata, James Clark and those who criticised XML Schema, pointed out the following:
>It is difficult to implement all features of XML Schema.
>It is difficult for engineers to read and write XML Schema definitions.
>It does not permit nondeterministic content models.
>Murata and collaborators designed another modern schema language, RELAX (Regular Language description for XML), more simple and mathematically consistent. They published RELAX specification in 2000. RELAX was approved as JIS and ISO/IEC standards. At roughly the same time, James Clark also designed another schema language, TREX (Tree Regular Expressions for XML).
>Murata and James Clark designed a new schema language RELAX NG based on TREX and RELAX Core. RELAX NG syntax is the expansion of TREX. RELAX NG was approved by OASIS in December 2001. RELAX NG was also approved as Part 2 of ISO/IEC 19757: Document Schema Definition Languages (DSDL).
https://en.wikipedia.org/wiki/Regular_Language_description_f...
https://en.wikipedia.org/wiki/RELAX_NG
https://en.wikipedia.org/wiki/XML_Schema_(W3C)
Schema Wars: XML Schema vs. RELAX NG (1/2) - exploring XML
https://web.archive.org/web/20180429143242/http://webreferen...
https://web.archive.org/web/20180429145711/http://webreferen...
https://news.ycombinator.com/item?id=22756875
>James Clark used Haskell to design and implement an algorithm for validating Relax NG XML schemas (he co-designed Relax NG, and designed its predecessor TREX), to work the ideas out before re-implementing it in (many many more lines of tedious brittle) Java (JING). Haskel works wonderfully as a design and standard definition language, that way.
https://news.ycombinator.com/item?id=25435678
>James Clark's compact syntax for Relax/NG XML schema validation language is quite tastefully designed, an equivalent but more convenient alternative syntax than XML, for writing tree regular expressions matching XML documents. It's way more beautiful and coherent than the official "XML Schema" standard.
[...]
>There's a wonderful DDJ interview with James Clark called "A Triumph of Simplicity: James Clark on Markup Languages and XML" where he explains how a standard has failed if everyone just uses the reference implementation, because the point of a standard is to be crisp and simple enough that many different implementations can interoperate perfectly.
>A Triumph of Simplicity: James Clark on Markup Languages and XML:
https://web.archive.org/web/20130721072712/https://www.drdob...
"The standard has to be sufficiently simple that it makes sense to have multiple implementations." -James Clark
<path>
<point x=10 y=10/>
<point x=90 y=10/>
<point x=90 y=90/>
<point x=10 y=90/>
<point x=10 y=10/>
</path>This could be a better format, but still a lot heaver in term of bandwidth.
<path>
<move x=10 y=10/>
<line x=90 y=10/>
<curve x=90 y=90/>
<end/>
</path> {
"path": [
[10, 10],
[90, 10],
[90, 90],
[10, 90],
[10, 10]
]
}
No need for the x and y, really. And it's still quite legible on one line. {"path": [[10, 10], [90, 10], [90, 90], [10, 90], [10, 10]]} {
paths: [
{
points: [
[10, 10], // 0
[90, 10], // 1
[90, 90], // 2
[10, 90] // 3
],
lines: [
[0, 1, 2, 3, 0] // connect point 0 to 1 to 2 to (...)
],
curves: {
0: [5, 15, 15, 5] // bezier control points for point 0 (x1, y1, x2, y2)
}
}
]
}
The downside (other than the bloat) would be when writing the code by hand, you'd have to keep track of the indices in the `points` array. [ { line: [10, 10] }, { cubic: [ 5, 15, 15, 5 ] }, ... ]
or even [ ["line", 10, 10], ["cubic", ...], ... ]
- anything as long as there's a way to stream the data (which is not possible with your format as you need to receive the whole object before being able to do anything).No, I didn't. I was working from the parent's code. I don't appreciate your attitude though. There's much nicer ways to correct people online.
> No need for the x and y, really
<path>
<m x=“10” y=“10” />
<h dx=“90” />
<v dy=“90” />
<h dx=“10” />
<l x=“10” y=“10” />
</path>
Perfectly sane.Also just in general, SVG's d-path syntax is a fantastic example of smart meeting of and understanding of requirements and users. If they used something more verbose (like `move-to` instead of `M` or whatever), then the file becomes WAY less human readable (too much noise at the XML-level), but more novice friendly. But... will novices need to edit d-paths manually? Of course not! Will they need to edit the XML structure manually? Way more likely. So you can use a more compact expert syntax for paths, because the people who will want to use it will likely be experts. It also means that SVG files are smaller, require less memory to parse, and more human-readable. Seriously great choice on their part.
>XML/SGML are a very effective way of representing tree data
Completely disagree:
- It's not effective density-wise because the format is very verbose
- It's not effective parsing-wise because the format is very complicated.
- It's not effective human-wise because you have meaningless distinctions between attributes and child nodes which makes sense for a markup language but not for a serialization format. It's also very verbose which makes it annoying to read and write.
Imagine that you have an object like:
struct Object {
name: "foo",
}
Should you serialize like: <object name="foo" />
Or: <object> <name>foo</name> </object>
And to be clear, that's a rhetorical question. My point is that this distinction doesn't make a lot of sense for a serialization format and forces the developer to make pointless decisions. Here you might say that an attribute makes more sense, but then if later you realize that you can have several names you're screwed, because you either have to come up with your own custom format to store those in an attribute (coma-separated? But what if there's a coma in the name?) or split them off into child nodes.It makes perfect sense for a markup language though, because there's (usually) a clear distinction between the textual and metatextual content. The stuff between the tags is meant to be displayed to the user, whereas the attributes are meant to be interpreted by the machine.
XML just tempts you to use attributes to represent strings, then kicks you in the ass when you least expect it by corrupting your data unexpectedly and unfairly but with full legal authority of the xml standard itself.
I wrote about it in this other comment:
(Note: XAML does have a nice duality between attributes and nodes (sort of) which is kind of interesting. It's been a while since I've used it though, so can't quite recall the details).
In XML ID fields are supposed to be unique to the document. Nobody enforces this. Instead getByID returns the first ID. Which means if you ask a DOM element and the root element for an element by ID, you can get two different answers.
Composing multiple schemas into the same file is a wordy, confusing mess, DTDs are straight up broken, and the only time I ever saw someone generate a Grade A XML Schema was by feeding examples into XMLSpy instead of writing their own.
XML asks a question we already knew the answer to: What if we made everyone into a programming language designer? And the answer was "anarchy" because we know that many people cannot design a syntactically consistent language, and hardly anybody can design a semantically consistent one.
James Strachan, as I recall, retired from Groovy before they ever figured out an unambiguous grammar for it. There comes a point where you realize you've made a mess that you are not qualified to clean up (Kernighan's Law). You can either do the very hard work of maturing into the responsibility, or bow out. Strachan is no different than a dozen people I've worked with and countless people I've heard about second or third hand, who I'm more than a little glad moved on.
The guilt is not theirs alone. Part must fall to this shared delusion that you can do anything with software if you only put your mind to it. We have mathematical proofs that tell us that's not true, and yet we still believe in the power of belief. Unfortunately if we didn't believe it a little, then we'd probably never write anything at all, so I'm not sure there's so much a cure as a condition that has to be managed. A little bit can go a long way, and most of us take it too far, sooner or later.
If I want something free-form text based, I've been more happy with e.g. YAML based formats.
<path d="M 10 10 H 90 V 90 H 10 L 10 10"/>
ctx.moveTo(10, 10);
ctx.lineTo(90, 10);
ctx.lineTo(90, 90);
ctx.lineTo(10, 90);
ctx.lineTo(10, 10);
Not much better. <draw id="my-path">
<move-to x="10" y="10" />
<line-to x="90" y="10" />
<line-to x="90" y="90" />
<line-to x="10" y="90" />
<line-to x="10" y="10" />
</draw>
<path draw="my-path" />
EDIT: Come to think of it, the above is probably a lot easier to animate using CSS or SMIL. Animating the `d` attribute on `<path>`s is theoretically possible but in practice weird at best or browser inconsistent or simply impossible at worst.EDIT 2: In D3 it is not uncommon to draw area graphs with a thick border on top. The only way to do this is with two `<path>`s with almost identical `d` attributes. Allowing drawings to be strung together with multiple idrefs could solve that:
<draw id=line-path><!-- ... --></draw>
<draw id="area-box-close">
<line-to x="100" y="100" />
<line-to x="0" y="100" />
<end />
</draw>
<path draw="line-path" class="line" />
<path draw="line-path area-box-close" class="area" />
Maybe this isn’t such a bad idea after all.I SVG is pretty inelegant, but it's pragmatical.
Also: Reminder SVGs goal is presentation, so paths even for a simple monochrome icon have _hundreds_ of control points. Something more intricate or a larger graphic? Thousands easily. All the examples here are toy examples, that don't scale to what SVG is actually used for.