Invisible XML is a language for describing the implicit structure of data
invisiblexml.org
invisiblexml.org
My only quandary would be whether the output XML structure could be ambiguous given the parse tree and input (requiring lots of context-dependent if/then logic when interpreting the XML). Perhaps some kind of invisible XML stylesheet could pre-process the AST before outputting the XML.
And secondly, can it handle CSV? If so, along with a command-line app like `jq` it could be an extremely useful addition to the general purpose data munging toolkit. Or do I have to pass the input through a 1000+ byte `sed` script first to normalize it.
I was literally just working on a project with TatSu[1] in Python which contains language elements that are conditional based on patterns in the syntax. I found that the added RegEx matching to EBNF-like syntax was quite powerful.
My only issue was the generated parser appeared to parse the entire content into memory rather than stream parsing.
I haven’t read enough on ixml yet, but, while it seems like while it would unlock many existing toolkits (I never hated XML but went away from it as the industry did), it seems like parsing through an XML format should be done in a stream mapping fashion, and not persisting data, for the XML to be truly “invisible”.
Adding the overhead of XML back into the processing chain… hmm, honestly have to think if the value of accessing the data and XML toolchain is worth it.
I’d almost rather see something that can read any input as a stream with a grammar and produce a stream (that can be materialized) of more optimized, yet open format, that can be compressed but handle complex types.. like protobufs, or flatbuffers. I’m not sure that humans need to read raw data files so long as we have great and open tooling to view data in binary formats (iirc, the biggest argument against XML and the added overhead originally).
Interesting although this seems a little out of date because xml seems to one of the least desirable data formats for modern programming languages. Maybe this is more useful for legacy enterprise use cases that I don’t know about.
This is essentially a way to write grammars for things and get the ability to parse them as trees in a common format that is interchangeable with things like JSON.
I honestly don't see much of a difference between this, and something like a PEG grammar where you do this:
let parsed = peg.parse(input, grammar)
let xml = json2xml(parsed)Ultimately what all that means is that all e.g. json documents can be represented losslessly in xml, whereas the reverse is not true without explicit external schema. Which means targeting XML covers other less capable interchange formats implicitly.
There is a real advantage of XML over JSON not mentioned though, which is its usefulness in annotating computer-readable data into an otherwise human-editable document. There's not a lot of these cases, and where they're at you're probably still better off using AsciiDoc, Markdown or even HTML instead, but those use cases are out there and JSON is awful for those.
Granted, JSON has achieved this same level of universality, but everything else (excepting perhaps CSV files) either suffers from obscurity or from weak/ambiguous/competing specifications.
The dream of semantically rich documents that XML provides (like, say, being able to cleanly interweave MathML, SVG, and XHTML in one document) is unmatched.
I think we'd see a lot more XML usage if it hasn't been over-promised, over-delivered (WS-*), and over-used (enterprise Java). If XML had stuck to its lane (making a schema language like RelaxNG instead of XmlSchema, for one thing) it wouldn't have left such a bad taste in so many people's mouths.
Yes, that's my point. If a human is required to look at the JSON it's not hard to find a parser that permits comments. So of course it would be supported if it was actually needed.
Even without that though, you can do something like the trick I used in a JSON-based format:
{ "__comment": "For help with this file, go to http://wiki.example.com/...", ... "foo": "..." }
That does have the problem that your programs processing the JSON need to strip out keys with that name if it would interfere with the program. So it's probably easier just to add the line of code in your parser saying to permit comments.
Meanwhile, XML comments have issues of their own, as you have to be familiar with SGML rules on how -- are handled to safely use or edit XML comments. So while they are nice to have it's not as if they have no nuance either.
Now, there is a trade off in more complex paths to interoperability with Uncle Joe and his DTD. It is still easier than trying to parse some convoluted json coming from yet another api that the dev team never dogfed a day in their life - because you can jump levels so precisely in xpath due to those verbose types and attributes. So as to XML verbosity at least, which for many appears to be the main complaint, for me it is worth it when done right.
Oh, here: https://www.rfc-editor.org/rfc/rfc8785
No duplicates, no whitespace, sort the keys, copy number serialization from javascript, a couple other little details.
This only two years old while JSON is in common usage for more than a decade... Without a spec there's dozen way to achieve it, for example by sorting keys using UTF-8 instead of UTF-16 values like done in this document, and the slightest difference would break things when used with crypto.
XSugar makes it possible to manage dual syntax for XML languages. An XSugar specification is built around a context-free grammar that unifies the two syntaxes of a language. Given such a specification, the XSugar tool can translate from alternative syntax to XML and vice versahttps://gist.github.com/felixjones/f8a06bd48f9da9a4539f
How do I implement a C++ library to parse PMX to XML and from that XML back into PMX?