JSON-schema is a thing as well, maybe not quite as mature (or "aged"), but the tooling is fine.
1. It's super verbose. Everybody hates writing it. And yes, people do write JSON by hand.
2. It has a weird and confusing data model with both attributes and child elements. Do you do <foo a="1"/> or <foo><a>1</a></foo>? In many many cases it's ambiguous and you end up with a weird mix, whereas in JSON it's obvious and easy.
https://en.wikipedia.org/wiki/RELAX_NG
RELAX NG has a clearly focused sound mathematical underpinning: regular expressions applied to trees, while XML Schema is an ad-hoc hot mess designed by committee.
https://en.wikipedia.org/wiki/XML_schema#RELAX_NG
Any decent JSON schema language should take its cues and learn from the design of RELAX NG, and not repeat the mistakes of XML Schema.
Makoto Murata raised some critical points about XML Schema, which was beyond repair, so they both invented new regexp-based schema languages to address those problems, and combined their respective work into Relax/NG:
https://en.wikipedia.org/wiki/Makoto_Murata#RELAX_and_RELAX_...
>Some people, including Murata and James Clark, had critical attitudes toward XML Schema. XML Schema is a modern XML schema language designed by W3C XML Schema Working Group. W3C intended XML Schema to supersede traditional DTD (Document Type Definition). XML Schema supports so many features that its specification is large and complex. Murata, James Clark and those who criticised XML Schema, pointed out the following:
>It is difficult to implement all features of XML Schema.
>It is difficult for engineers to read and write XML Schema definitions.
>It does not permit nondeterministic content models.
James Clark compared Relax NG to XML Schema and its predecessor, SGML Document Type Definitions, in his paper, "The Design of RELAX NG":
https://relaxng.org/jclark/design.html
James Clark has a huge amount of experience designing and implementing SGML and XML standards:
https://en.wikipedia.org/wiki/James_Clark_(programmer)
I've written about XML Schema, Relax NG, and James Clark eariler:
https://news.ycombinator.com/item?id=26122033
>Some of the most incredibly awfully bad XML DSLs are official standards, themselves. COUGH XMLSchema COUGH: [...]
https://news.ycombinator.com/item?id=22756875
>James Clark used Haskell to design and implement an algorithm for validating Relax NG XML schemas (he co-designed Relax NG, and designed its predecessor TREX), to work the ideas out before re-implementing it in (many many more lines of tedious brittle) Java (JING). Haskel works wonderfully as a design and standard definition language, that way. [...]
A Triumph of Simplicity: James Clark on Markup Languages and XML
https://www.drdobbs.com/a-triumph-of-simplicity-james-clark-...
>If you peek under the hood of high-profile open-source projects such as Mozilla, Apache, Perl, and Python, you'll find a little program called "expat" handling the XML parsing. If you've ever used the man command on your GNU/Linux distribution, then you've also used groff, the GNU version of the UNIX text formatting application, troff. If you've ever done any work with SGML, from generating documentation from DocBook to building your own SGML applications, you've undoubtedly come across sgmls, SP, and Jade.
This article "Schema Wars: XML Schema vs. Relax NG (1/2) - exploring XML" enumerates some of James Clark's objections to XML Schema:
https://web.archive.org/web/20190403171512/http://webreferen...
>James Clark, leader of the technical committee at OASIS for RELAX NG, and author of one of the first XML parsers, recently described the problems of the XML Schema language in a newsgroup posting:
>1) XML Schema definitions require considerable expertise to understand and can contain quite a few surprises.
>As an example, if you derive a complex type by restriction you have to specify the new restricted content model explicitly. However, attributes are treated in the opposite way: by default you get all the attributes and you have to explicitly rule out the ones you don't want. A similar inconsistency exists in that if you merge two attribute definitions you get the union of the concrete attribute definitions, but the intersection of attribute wildcards (specified by the anyAttribute element). While these might be convenient choices for the specification, it easily creates confusion for the human reader of XML schemas.
>2) The XML Schema Recommendation is hard to read and understand.
>To avoid the possible misinterpretations mentioned above you might have to reference the specification in order to fully understand a specific schema definition. I have to agree with Clark that the W3C's XML Schema Recommendation is by far the hardest to read and understand, making it even more difficult to make sense of a particular schema presented to you.
>3) W3C XML Schema's support for attributes provides no advance over DTDs.
>As with DTDs, W3C XML Schema only allows the specification of whether attributes are required or optional. There is no way to specify more complex constraints between attributes or between attributes or elements, for instance that either attribute X or attribute Y is allowed or that either attribute X or element Y is allowed. The mechanism that is used to constrain the co-occurrence of child elements should be extended to attributes and the combinations of attributes and child elements.
>4) W3C XML Schema provides very weak support for unordered content.
>When the designer of an XML vocabulary does not wish to force child elements to occur in a particular order, it can be impractical to describe the XML vocabulary using XML Schema, because XML Schema imposes such limitations.
>5) Datatype handling in W3C XML Schema lacks modularity.
>W3C XML Schema is tied to the single collection of datatypes defined in Part 2 of W3C XML Schema. Yet this collection of datatypes is a very ad-hoc collection. It includes datatypes of highly debatable relevance (gYearMonth, gDay etc). Yet it lacks many datatypes that are important for many applications. A modular approach where a schema language can be combined with one or more standard collections of datatypes, some general-purpose and some domain-specific, is called for here.
>6) W3C XML Schema does not define a single notion of validity of a document with respect to a schema.
>There are different varieties of validation (lax and strict) and many different ways to validate a document against a schema. From a W3C XML Schema alone, it is not possible to know what is a valid document. For instance there is no way to specify what is allowed as the root element.
>7) Magic schema attributes in documents
>W3C XML Schema provides the xsi:schemaLocation attribute, which allows an XML document instance to indicate the schema that should be used to validate the document. This creates problems with security (the destination might have changed or tampered with), interoperability (use of schemaLocation is optional) and "purity" of schema definition: There is no way to prevent the document containing magic xsi:* attributes, so the use of W3C XML Schema "infects" the grammar you are defining.
>8) Another problematic area in W3C XML Schema is the support for infoset augmentation, such as default attributes.
>Apart from being a violation of modularity, this tends to cause interoperability problems, because it leads to the possibility of the application getting different information depending on whether or not validation has been performed.
For example, you might want to perform one kind of quick simple input sanity checking validation that ensures the input won't crash your processor, to perform in real time on millions of XML documents you receive every minute.
Then you might also want another more elaborate schema with an anal retentive formal specification, including text syntax checking, data type validation, cross-reference checking, constraints between values, and documentation, for creating and editing the same document in a schema-driven editor.
https://github.com/michmech/xonomy
It's also a bad idea to interpret the URI of a schema as the URL from which you can download the schema definition. That introduce security and privacy and interoperability problems.
Internet Explorer used to try to download the schema definition every time you opened an XML document, which lets whoever controls that web site know you're looking at one of their documents, and what its url is in the referer field. It's like XML's answer to the 1x1 transparent GIF tracking cookie!
Unlike DTDs and XML Schema, RELAX NG doesn't presume that it's the only schema language in the world that you will ever need to use: there are other useful languages like Schematron, which do things that RELEX NG was designed not to do on purpose.
(There's a great Ven diagram comparing XSD, DTD, RELAX NG and Schematron's scope on that page.)
https://www.schematron.com/document/2755.html
>Schematron can be useful in conjunction with many grammar-based structure-validation languages: DTDs, XML Schemas, RELAX, TREX, etc. Indeed, Schematron is part of an ISO standard (DSDL: Document Schema Description Languages) designed to allow multiple, well-focussed XML validation languages to work together. You can even embed a Schematron schema inside an XML Schema <appinfo> element or inside a RELAX NG schema!
James Clark on RELAX NG and xml:id:
https://blog.jclark.com/2009/01/relax-ng-and-xmlid.html
>RELAX NG and xml:id
>One part of the vision underlying RELAX NG is that validation should not be monolithic: it is not necessary or desirable to have one schema language that can handle every possible kind of validation you might want to do; it is better instead to have multiple specialized languages, each of which does one kind validation, really well. Consistent with this vision, RELAX NG provides only grammar-based validation. There's no implicit claim that other kinds of validation aren't useful and important.
>One kind of validation that is clearly useful and important and that can't be done by grammars is checking of cross-references. One possibility is to use Schematron for this. The designers of RELAX NG anticipated that there would be a little schema language specialized to this, which would be created as part of the ISO DSDL effort (as part 6); this wouldn't be a million miles from the kind of thing that XSD provides with xs:key/xs:unique/xs:keyref. Unfortunately this hasn't happened yet.
[...]
It's not all goofy and arbitrary and complex line noise like XML DTDs, and it's not all verbose and clumsy namespaced XML tags and attributes like XML Scheme documents.
It even supports inheritance, and grammar as well as element level annotations (useful for schema and meta processing tools and editors).
RELAX NG's Compact Syntax:
https://www.xml.com/pub/a/2002/06/19/rng-compact.html
RELAX NG Compact Syntax: Committee Specification 21 November 2002:
https://www.oasis-open.org/committees/relax-ng/compact-20021...
RELAX NG Compact Syntax Tutorial:
Like, how do you pronounce that? I want to say 'Relax-ung', but it sounds odd, like you're trying to say "relaxing" but failing; or maybe 'Relax Nig', but that sounds like something you'd see in an HN thread on Timnit Gebru's firing. It doesn't feel like the naming was considered from a normal human point of view, as opposed to from the golden mountain of abstraction.
They could have pulled a slick smooth move and called it "TREX-LAX", for "Tree Regular EXpressions for LAnguages of Xml", which rolls off the tongue more mellifluously!
TREX-LAX: It makes your XML nice and regular, and keeps you shit moving smoothly.
https://www.browncafe.com/community/threads/fedex-and-ex-lax...
>FedEx and Ex-Lax battle for rights to promotional slogan
If you want to drop JSON altogether, something like Protobufs would be better. But JSON + JSONSchema has all the benefits of XML and none of the drawbacks.
XML Schema were supposed to fix all the problems of DTDs, but they don't support unordered content, and have a host of their own problems. XML Schema are typically external XML documents referenced by an xsi:schemaLocation attribute, but apparently you can embed them too. But I don't know if that's a common practice, since it's a pretty bad idea for a lot of reasons:
https://www.herongyang.com/XML/XSD-Statements-Embedded-in-XM...
You raise a good point that it's a bad idea to embed schemas in documents, or even references to schemas in documents, because, as James Clark points out, magic schema attributes in documents (like xsi:schemaLocation) pre-suppose that there can be only one way of validating a document, they create security and interoperability problems, and it infects documents with the namespace of the grammar that you're using to validate them, when documents shouldn't depended on such knowledge, and be forced to include hot messes of namespace attributes like:
<beans xmlns="http://www.springframework.org/schema/beans" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:aop="http://www.springframework.org/schema/aop" xsi:schemaLocation="http://www.springframework.org/schema/beans http://www.springframework.org/schema/beans/spring-beans-2.0... http://www.springframework.org/schema/aop http://www.springframework.org/schema/aop/spring-aop-2.0.xsd">
Schema Wars: XML Schema vs. RELAX NG (1/2) - exploring XML:
https://web.archive.org/web/20190403171512/http://webreferen...
>James Clark, leader of the technical committee at OASIS for RELAX NG, and author of one of the first XML parsers, recently described the problems of the XML Schema language in a newsgroup posting: [...]
>7) Magic schema attributes in documents
>W3C XML Schema provides the xsi:schemaLocation attribute, which allows an XML document instance to indicate the schema that should be used to validate the document. This creates problems with security (the destination might have changed or tampered with), interoperability (use of schemaLocation is optional) and "purity" of schema definition: There is no way to prevent the document containing magic xsi:* attributes, so the use of W3C XML Schema "infects" the grammar you are defining.
Here's more of what he has to say about JSON in his blog post "XML vs the Web":
https://blog.jclark.com/2010/11/xml-vs-web_24.html
>[...] From this perspective, my reaction to JSON is a combination of "Yay" and "Sigh".
>It's "Yay", because for important use cases JSON is dramatically better than XML. In particular, JSON shines as a programming language-independent representation of typical programming language data structures. This is an incredibly important use case and it would be hard to overstate how appallingly bad XML is for this. The fundamental problem is the mismatch between programming language data structures and the XML element/attribute data model of elements. This leaves the developer with three choices, all unappetising:
>- live with an inconvenient element/attribute representation of the data;
>- descend into XML Schema hell in the company of your favourite data binding tool;
>- write reams of code to convert the XML into a convenient data structure.
>By contrast with JSON, especially with a dynamic programming language, you can get a reasonable in-memory representation just by calling a library function. [...]
>There's a bigger point that I want to make here, and it's about the relationship between XML and the Web. When we started out doing XML, a big part of the vision was about bridging the gap from the SGML world (complex, sophisticated, partly academic, partly big enterprise) to the Web, about making the value that we saw in SGML accessible to a broader audience by cutting out all the cruft. In the beginning XML did succeed in this respect. But this vision seems to have been lost sight of over time to the point where there's a gulf between the XML community and the broader Web developer community; all the stuff that's been piled on top of XML, together with the huge advances in the Web world in HTML5, JSON and JavaScript, have combined to make XML be perceived as an overly complex, enterprisey technology, which doesn't bring any value to the average Web developer.
>This is not a good thing for either community (and it's why part of my reaction to JSON is "Sigh"). XML misses out by not having the innovation, enthusiasm and traction that the Web developer community brings with it, and the Web developer community misses out by not being able to take advantage of the powerful and convenient technologies that have been built on top of XML over the last decade.
>So what's the way forward? I think the Web community has spoken, and it's clear that what it wants is HTML5, JavaScript and JSON. XML isn't going away but I see it being less and less a Web technology; it won't be something that you send over the wire on the public Web, but just one of many technologies that are used on the server to manage and generate what you do send over the wire.
>In the short-term, I think the challenge is how to make HTML5 play more nicely with XML. In the longer term, I think the challenge is how to use our collective experience from building the XML stack to create technologies that work natively with HTML, JSON and JavaScript, and that bring to the broader Web developer community some of the good aspects of the modern XML development experience.