https://en.wikipedia.org/wiki/RELAX_NG
RELAX NG has a clearly focused sound mathematical underpinning: regular expressions applied to trees, while XML Schema is an ad-hoc hot mess designed by committee.
https://en.wikipedia.org/wiki/XML_schema#RELAX_NG
Any decent JSON schema language should take its cues and learn from the design of RELAX NG, and not repeat the mistakes of XML Schema.
Makoto Murata raised some critical points about XML Schema, which was beyond repair, so they both invented new regexp-based schema languages to address those problems, and combined their respective work into Relax/NG:
https://en.wikipedia.org/wiki/Makoto_Murata#RELAX_and_RELAX_...
>Some people, including Murata and James Clark, had critical attitudes toward XML Schema. XML Schema is a modern XML schema language designed by W3C XML Schema Working Group. W3C intended XML Schema to supersede traditional DTD (Document Type Definition). XML Schema supports so many features that its specification is large and complex. Murata, James Clark and those who criticised XML Schema, pointed out the following:
>It is difficult to implement all features of XML Schema.
>It is difficult for engineers to read and write XML Schema definitions.
>It does not permit nondeterministic content models.
James Clark compared Relax NG to XML Schema and its predecessor, SGML Document Type Definitions, in his paper, "The Design of RELAX NG":
https://relaxng.org/jclark/design.html
James Clark has a huge amount of experience designing and implementing SGML and XML standards:
https://en.wikipedia.org/wiki/James_Clark_(programmer)
I've written about XML Schema, Relax NG, and James Clark eariler:
https://news.ycombinator.com/item?id=26122033
>Some of the most incredibly awfully bad XML DSLs are official standards, themselves. COUGH XMLSchema COUGH: [...]
https://news.ycombinator.com/item?id=22756875
>James Clark used Haskell to design and implement an algorithm for validating Relax NG XML schemas (he co-designed Relax NG, and designed its predecessor TREX), to work the ideas out before re-implementing it in (many many more lines of tedious brittle) Java (JING). Haskel works wonderfully as a design and standard definition language, that way. [...]
A Triumph of Simplicity: James Clark on Markup Languages and XML
https://www.drdobbs.com/a-triumph-of-simplicity-james-clark-...
>If you peek under the hood of high-profile open-source projects such as Mozilla, Apache, Perl, and Python, you'll find a little program called "expat" handling the XML parsing. If you've ever used the man command on your GNU/Linux distribution, then you've also used groff, the GNU version of the UNIX text formatting application, troff. If you've ever done any work with SGML, from generating documentation from DocBook to building your own SGML applications, you've undoubtedly come across sgmls, SP, and Jade.
This article "Schema Wars: XML Schema vs. Relax NG (1/2) - exploring XML" enumerates some of James Clark's objections to XML Schema:
https://web.archive.org/web/20190403171512/http://webreferen...
>James Clark, leader of the technical committee at OASIS for RELAX NG, and author of one of the first XML parsers, recently described the problems of the XML Schema language in a newsgroup posting:
>1) XML Schema definitions require considerable expertise to understand and can contain quite a few surprises.
>As an example, if you derive a complex type by restriction you have to specify the new restricted content model explicitly. However, attributes are treated in the opposite way: by default you get all the attributes and you have to explicitly rule out the ones you don't want. A similar inconsistency exists in that if you merge two attribute definitions you get the union of the concrete attribute definitions, but the intersection of attribute wildcards (specified by the anyAttribute element). While these might be convenient choices for the specification, it easily creates confusion for the human reader of XML schemas.
>2) The XML Schema Recommendation is hard to read and understand.
>To avoid the possible misinterpretations mentioned above you might have to reference the specification in order to fully understand a specific schema definition. I have to agree with Clark that the W3C's XML Schema Recommendation is by far the hardest to read and understand, making it even more difficult to make sense of a particular schema presented to you.
>3) W3C XML Schema's support for attributes provides no advance over DTDs.
>As with DTDs, W3C XML Schema only allows the specification of whether attributes are required or optional. There is no way to specify more complex constraints between attributes or between attributes or elements, for instance that either attribute X or attribute Y is allowed or that either attribute X or element Y is allowed. The mechanism that is used to constrain the co-occurrence of child elements should be extended to attributes and the combinations of attributes and child elements.
>4) W3C XML Schema provides very weak support for unordered content.
>When the designer of an XML vocabulary does not wish to force child elements to occur in a particular order, it can be impractical to describe the XML vocabulary using XML Schema, because XML Schema imposes such limitations.
>5) Datatype handling in W3C XML Schema lacks modularity.
>W3C XML Schema is tied to the single collection of datatypes defined in Part 2 of W3C XML Schema. Yet this collection of datatypes is a very ad-hoc collection. It includes datatypes of highly debatable relevance (gYearMonth, gDay etc). Yet it lacks many datatypes that are important for many applications. A modular approach where a schema language can be combined with one or more standard collections of datatypes, some general-purpose and some domain-specific, is called for here.
>6) W3C XML Schema does not define a single notion of validity of a document with respect to a schema.
>There are different varieties of validation (lax and strict) and many different ways to validate a document against a schema. From a W3C XML Schema alone, it is not possible to know what is a valid document. For instance there is no way to specify what is allowed as the root element.
>7) Magic schema attributes in documents
>W3C XML Schema provides the xsi:schemaLocation attribute, which allows an XML document instance to indicate the schema that should be used to validate the document. This creates problems with security (the destination might have changed or tampered with), interoperability (use of schemaLocation is optional) and "purity" of schema definition: There is no way to prevent the document containing magic xsi:* attributes, so the use of W3C XML Schema "infects" the grammar you are defining.
>8) Another problematic area in W3C XML Schema is the support for infoset augmentation, such as default attributes.
>Apart from being a violation of modularity, this tends to cause interoperability problems, because it leads to the possibility of the application getting different information depending on whether or not validation has been performed.
For example, you might want to perform one kind of quick simple input sanity checking validation that ensures the input won't crash your processor, to perform in real time on millions of XML documents you receive every minute.
Then you might also want another more elaborate schema with an anal retentive formal specification, including text syntax checking, data type validation, cross-reference checking, constraints between values, and documentation, for creating and editing the same document in a schema-driven editor.
https://github.com/michmech/xonomy
It's also a bad idea to interpret the URI of a schema as the URL from which you can download the schema definition. That introduce security and privacy and interoperability problems.
Internet Explorer used to try to download the schema definition every time you opened an XML document, which lets whoever controls that web site know you're looking at one of their documents, and what its url is in the referer field. It's like XML's answer to the 1x1 transparent GIF tracking cookie!
Unlike DTDs and XML Schema, RELAX NG doesn't presume that it's the only schema language in the world that you will ever need to use: there are other useful languages like Schematron, which do things that RELEX NG was designed not to do on purpose.
(There's a great Ven diagram comparing XSD, DTD, RELAX NG and Schematron's scope on that page.)
https://www.schematron.com/document/2755.html
>Schematron can be useful in conjunction with many grammar-based structure-validation languages: DTDs, XML Schemas, RELAX, TREX, etc. Indeed, Schematron is part of an ISO standard (DSDL: Document Schema Description Languages) designed to allow multiple, well-focussed XML validation languages to work together. You can even embed a Schematron schema inside an XML Schema <appinfo> element or inside a RELAX NG schema!
James Clark on RELAX NG and xml:id:
https://blog.jclark.com/2009/01/relax-ng-and-xmlid.html
>RELAX NG and xml:id
>One part of the vision underlying RELAX NG is that validation should not be monolithic: it is not necessary or desirable to have one schema language that can handle every possible kind of validation you might want to do; it is better instead to have multiple specialized languages, each of which does one kind validation, really well. Consistent with this vision, RELAX NG provides only grammar-based validation. There's no implicit claim that other kinds of validation aren't useful and important.
>One kind of validation that is clearly useful and important and that can't be done by grammars is checking of cross-references. One possibility is to use Schematron for this. The designers of RELAX NG anticipated that there would be a little schema language specialized to this, which would be created as part of the ISO DSDL effort (as part 6); this wouldn't be a million miles from the kind of thing that XSD provides with xs:key/xs:unique/xs:keyref. Unfortunately this hasn't happened yet.
[...]
It's not all goofy and arbitrary and complex line noise like XML DTDs, and it's not all verbose and clumsy namespaced XML tags and attributes like XML Scheme documents.
It even supports inheritance, and grammar as well as element level annotations (useful for schema and meta processing tools and editors).
RELAX NG's Compact Syntax:
https://www.xml.com/pub/a/2002/06/19/rng-compact.html
RELAX NG Compact Syntax: Committee Specification 21 November 2002:
https://www.oasis-open.org/committees/relax-ng/compact-20021...
RELAX NG Compact Syntax Tutorial:
Like, how do you pronounce that? I want to say 'Relax-ung', but it sounds odd, like you're trying to say "relaxing" but failing; or maybe 'Relax Nig', but that sounds like something you'd see in an HN thread on Timnit Gebru's firing. It doesn't feel like the naming was considered from a normal human point of view, as opposed to from the golden mountain of abstraction.
They could have pulled a slick smooth move and called it "TREX-LAX", for "Tree Regular EXpressions for LAnguages of Xml", which rolls off the tongue more mellifluously!
TREX-LAX: It makes your XML nice and regular, and keeps you shit moving smoothly.
https://www.browncafe.com/community/threads/fedex-and-ex-lax...
>FedEx and Ex-Lax battle for rights to promotional slogan