A Journey building a fast JSON parser and full JSONPath
github.com
github.com
I’m going to invent Baz’s 11th law of computing here: any data format that isn’t XML will evolve into a badly specified version of XML over time.
For machine-to-machine communication, it's very well suited, but most data is simple enough, and the XML libraries I've used tended to be --let's say-- over-engineered, while there are no hoops to jump through when you want to parse JSON.
And one thing I always disliked about XML was the CDATA section: it makes the message even harder to read, and it's not like you're going to use that binary data unparsed/unchecked.
XML just tried to formalize data transfer and description prematurely, which made it rigid and not even sufficiently powerful. I must say that XSLT and XPath were great additions, though.
Compare that to XML where we have a plethora of established tools (Woodstoxx, JAXB, etc.).
What I have trouble to understand, which everybody else just seems to accept as obvious, is why one would take on these problems? Is JSON Schema more powerful than XML Schema? Does the use of JSON have advantages over using XML? When we are talking about a client program calling a server API with JSON/XML, why do we care about the format of data exchanged? What advantages does JSON have in this case in contrast to XML (or for that matter a binary format like Protocol Buffers)? Isn't this the most boring part of the application, which you would want to just get out of the way and work? What are the advantages of JSON over XML that would lead me to deal with the problems of evolving specifications and unreliable tooling?
(And just to repeat, since everybody seems to have a different opinion about this than me, I must be missing something and really would like to learn what!)
OpenAPI is probably used a bit more than json schema, but it's contextually limited to APIs (which, to be fair, is mostly what JSON is used for).
OpenAPI is another example. There are threads on hacker news about generating code from OpenAPI specs. These always seem to say "oh, yes don't use tool X, use tool Y it does not have that problem, although it also doesn't support Z". The consensus seems to be to not generate code from an OpenAPI specification but to just use it as documentation, since all generators are more or less broken. Contrast that with for example JAXB (which is not an exact replacement I know), which has been battle tested for years.
>The consensus seems to be to not generate code from an OpenAPI specification but to just use it as documentation, since all generators are more or less broken.
OpenAPI still functions just fine as a means of documentation and validation.
I'm allergic to all forms of code generation, to be honest. If there is an equivalent of XML in this I imagine it's even more horrendous. I can just imagine chasing down compiler errors indirectly caused by an XML switch not set shudder.
>Contrast that with for example JAXB
JAXB looks like a bolt on to work around XML's deficiencies. There's no need to marshal JSON to special funky data structures in your code because lists and hashmaps are already built in. You can just use those. An equivalent doesn't need to exist.
For schema validation, I think XML has, what, 3 ways of doing it? DTDs? XMLSchema? And now JAXB does a bit of that on the side too? Does that sound like a healthy ecosystem to you? Because it sounds like absolute dogshit to me.
WSDL comes to mind
Honestly the same issue with versioning has been my primary issue with XML Schemas in the past. XSD 1.1 for example came out over a decade ago, but is still very badly supported in most tooling I tried out.
> When we are talking about a client program calling a server API with JSON/XML, why do we care about the format of data exchanged?
We shouldn't care much, beyond debuggability (can a developer easily see what's going on), (de)serialization speed, and bandwith use. JSON and protobuf tend to be a decent chunk smaller than XML, JSON is a bit easier to read, and Protobuf is faster to (de)serialize. This means they should generally be preferred.
In the case of a client program calling a server API I'd personally have the server do the required validation on a deserialized object, instead of doing so through a schema. This is generally easier to work on for all developers in my team, and gets around all the issues with tooling. The only real reason I use schemas is when I'm writing a file by hand, and want autocompletion and basic validations. In that case versioning and tooling issues are completely in my control.
One of the first XSLT transforms I was ever given to maintain generated XML by the same method. <xsl:text><PRICE></xsl:text><xsl:value-of select="PRICE"/><xsl:text></PRICE></xsl:text> and so on.
That's why it never caught on.
The ability of JSON/Javascript to tape together kinda-working solutions before and instead of any kind of specification works is hugely powerful, because it allows iterating on the requirements by having actual users use the app.
I lived through all that, and can totally understand why people turned away in disgust and agreed on REST instead.
S-expressions are the most direct representation of a tree: (root node node ...). Trees are everywhere, they represent any nested structure; lists are logically a subset of trees.
XML is a tree. It has the weird "attribute" node types, a legacy of SGML text markup notation. JSON is a tree, obviously. So is protobuf, thrift, etc. They all could be serialized as s-expressions.
Now, a schema that destined a tree is also a tree. Hence XML Schema, JSONSchema, etc.
More, an abstract program that describes a transformation of a tree is also a tree; this products homoiconic languages, from XSLT to Lisps.
There is nothing special about XML; it's just a particular case of a generic law.
I've not looked at KDL before but a quick scan suggests it's interesting. I will look into it.
When you read the s-expression alternatives proposed to XML with an eye to "How would I actually code against this? How would I actually convince multiple people to use the exact same standard as me? How do I support all the use cases of interest to me?" they completely fall apart. They're too simple. The very fact I have to use the plural for s-expression alternative since no two of them are every quite the same says quite a bit.
When you need that structure, XML is actually a very good choice; the error people made was using it when they didn't need that structure. Note how much of the complaint about using XML, even in this very conversation, is (quite correctly!) "what do I do with all these extra structural elements?" If you don't have a clear answer to that, don't use XML. If you do, don't jam it into s-exprs or JSON either, you end up with an even worse mess.
The same way you agree on an XML schema? I don't know if I quite understand what you want to say - as I see it both are tree structured formats which means they both can represent the same information, just that s-expressions are less verbose but XML has more existing tooling for defining & validating a structure. Though the latter is more an aspect of the ecosystem than the format itself.
Yes.
But my point is, there's a lot of people who seem to think just waving the word "s-expression" at the problem is a solution. If you have to sit down and carefully define the exact standard, the value over XML gets mitigated a lot, because the standard is going to need a lot of stuff in it. It has to at least be as complicated as JSON, and JSON is often not quite enough. You can't just say "why don't we use s-expressions" as if that's an argument on its own.
But the reality is JSON ate that space anyhow. There's no room for something that's "like JSON, but not really" anymore.
XML’s decline from its peak of adoption mean lots of people working with data disagree with you.
Every data format will eventually evolve into a tree
That more ancient law would Greenspun's tenth rule, FYI—or a corollary to it, at least.
The law proposed here (as Baz's 11th law) was intended to be a humorous and obvious pastiche crafted with Greenspun's quip in mind, with the idea being that the reader would be in on the joke (being already familiar with it).
So considering XML is subsetted from SGML, I guess the answer is closer to yes than thought.
Though probably it's worth citing the following quote from that paper:
> If the sweet spot for XML and SGML is marking up “prose documents”, the sweet spot for JSON is collections of atomic values.
[1]: https://www.balisage.net/Proceedings/vol17/html/Walsh01/Bali...
In the other words, SGML was way too powerful than what we actually needed. Of course we are with the benefit of hindsight though.
The widespread use of markdown and other lightweight markup rather than rigid XML-style fully tagged markup for authoring tells otherwise though. And so does the continued use of HTML chock full of SGMLisms such as tag inference and attribute shortforms that weren't included in the XML subset/profile when XML (XHTML) was created to replace HTML.
So while XML isn't used as an authoring format on the web (nor as delivery format), it's still useful as canonical archival format I guess.
<Foo>
<Shininess>HIGH</Shininess>
<Luck>7</Luck>
</Foo>
or <Foo shininess="HIGH" luck="7" />
and yet countless thousands decided to do just that, for reasons that are totally inexplicable to me.Obviously as a markup language this is fine; as a data format it's bizarre, since the division between attribute vs child doesn't match most in-memory data structures.
With some niche exceptions where it has clung on, XML basically died. It's time to move on. The fact that we do similar sorts of stuff with JSON like data transformations and schema validation does not, in any way, shape or form, invalidate its flaws.
The overcomplicated mess was the WS-* garbage.
no user knows
The described problem literally doesn't exist in XML. Your XML-validating editor will check your document against the schema and will not allow for an attribute where the sub-element is required and vice versa.JSON became only popular because of similar opportunistic effects (ie being already part of the stack via eval()). If you look at how typical non-JS backends such as Java or .net deal with service request/response data, there's absolutely no advantage for either JSON or XML - both are represented as class/structure and (de-)serialized via binding frameworks and annotations.
I guess you could argue we should all use Protocol Buffers, pickle, Thrift, etc.[1] and only switch to JSON for debugging. I wouldn't disagree. Protobuf is apparently faster than JSON in the browser.
[0] See https://www.interfaceware.com/hl7-message-structure for an example message
[1] I missed Corba and spent the early years of my professional life trying not to touch the SOAP, just in case I dropped it
vtables are attributes for pointers. hypergraphs (as used in some tagging systems) have attributes on everything, including attributes. CBOR has optional type-tags on its items.
Everyone is just confused because people who didn't know this designed HTML. But also everyone is confused because HTML and XML aren't necessarily related other than some parentage in SGML.
If anything, what's wrong with HTML in this respect is that JavaScript and CSS can be put inline into content when these should always go into attributes and/or external resources linked via src/href attributes. And this flaw shows indeed where HTML deviates from SGML proper: when the style and script elements were introduced, their "content" needed to be put into SGML comment tags <!-- and --> such that browsers wouldn't render JavaScript snd CSS as text content. I mean, who came up with this brain-dead design?
But CSS is a lost cause anyway. What does it tell you about its designers that they thought, starting with a markup language having pretty intense syntactic constructs already, to tunnel yet another item=value syntax in regular markup attributes? Like replacing <h2 bgcolor=black> by <h2 style="background-color: black"> and then claiming attributes are for "behavior" or whatever nonsense after the fact. Whoever came up with this clearly wasn't a CompSci person. And the syntactic proliferation in CSS completely became out of hand, for the simple reason that HTML evolution was locked down while W3C was focussed on XML/XHTML for over a decade, while the CSS spec process was lenient.
I suspect the popularity was due to the sax parser and “interop” between C++ and Java.
To me coming from ObjC++, json is just a serialised dictionary.
<key>..</key>
<value>..<value>
<key>..</key>
<value>..<value>
...
Taking something that supports explicit structure and making it purely ordering based.And it turns out that both JSON and XML are used for data interchange, and when people have data interchange problems, they build tooling to help solve those problems (like schema validation). That doesn't make JSON "like XML", it just means they're discovering the same problem and solving it for the format they're using.
How about doing comparisons against other implementations?
Like this one: https://github.com/json-iterator/go
Update: found this outdated repo: https://github.com/ohler55/compare-go-json
> Note that while JSON Pointer (RFC 6901) is already standardised, it is designed to provide a reference to a single, specific part of a JSON document, whereas JSONPath provides the ability to query a document and potentially return multiple values.
I made something similar in Java - unify-jdocs - https://github.com/americanexpress/unify-jdocs - though this is not for parsing - it is more for reading and writing when the structure of the document is known - read and write any JSONPath in one line of code and use model documents to define the structure of the data document (instead of using JSONSchema which I found very unwieldy to use) - no POJOs or model classes - along with many other features. Posting here as the topic is relevant and it may help people in the Java world. We have used it intensively within Amex for a very large complex project and it has worked great for us.
" A valid example of a SEN document is:
{ one: 1 two: 2 array: [a b c] yes: true } "
{:one 1 :two 2 :array [a b c] :yes true}
cf. https://github.com/edn-format/edn
Likewise, commas are considered whitespace. They are sometimes added to make lengthy maps easier to read.
{
"one": 1,
"two": 2,
"array": ["a", "b", "c"],
"yes": true
}
That example also caught my attention, but in a bad way. It looks just like a comeback of one of the worst ideas of YAML.My immediate question would be what's the JSON for this SEN I've crafted:
{
array: [string1 string2 "true" true True TRUE yes y]
}
For more fun, there's a single problematic entry here, can you spot it?: 1.20.4
1.204.4
1.20
1.204
1.20.0
1.20.00
1.20-rc2
Or, level expert, there's exactly one problem here as well: 0a1f
0bfd
0c0c
0d01
0e02I don't get your other examples, can you explain? I assumed that 1.20.4 is not a valid SEN entry, because it starts with a digit but is not a number.
The true/"true" leaves a bad taste in my mouth after yaml, but overall SEN is an improvement :)
` { array: ["string1" "string2" "true" true "True" "TRUE" "yes" "y"] } `
"Strings can also be delimited with a single quote character which allows for a string to be either "abc" or 'abc'."
There's no mention of having a string without a delimeter.
> array: [a b c]