- Support for Maps and Sets (unlike objects, maps allow arbitrary types to be used as keys)
- A standard Date format
- Embedded binary blobs. No idea how to do this and keep it human readable, but when you need this its super useful. Maybe something similar to WS's binary message encoding.
- Arbitrary precision integer support. This is particularly useful for cryptocurrencies and for interoperability with 64 bit integers in other languages. And bigints are coming to javascript - https://github.com/tc39/proposal-bigint
- Maybe even fix JSON's weird unicode encoding: http://timelessrepo.com/json-isnt-a-javascript-subset
Unfortunately it seems like nobody 'owns' JSON enough to give JSON 2.0 the political weight it would need for cross-language support.
$ node
> JSON.parse('1231231231231231231231231123123123123')
1.2312312312312312e+36
I've heard of JSON implementations in other languages making bigger JSON numbers than javascript supports; but I didn't realise JSON.parse would quietly parse them and throw away precision in the process. I wonder what the plan is for JSON serialization support of bigints. I assume there is no plan - Chrome stable 67 supports bigints, but JSON.serialize(2n) throws a TypeError.Which means, in practice, you are stuck passing numbers as strings (and on the javascript side, pass that string into the arbitrary precision number library)
And really, if you have consumers/producers in different languages/codebases, built-in support for these things is a great convenience. You wouldn't say that objects/dicts in JSON are irrelevant, would you?
I kind of would. Coming from Erlang, I don't see a point to there being a whole separate syntax for encoding object/map types. Just encode objects/maps as arrays of pairs (2-Tuples, or just length-2 arrays).
Then, if you want ser-des type fidelity, stick an annotation onto the array (like in YAML) to say what type it should decode to.
E.g., something like:
[1, 2, 3] # array
[['a', 1], ['b', 2], ['c', 3]] # array of pairs
[@object, ['a', 1], ['b', 2], ['c', 3]] # array of pairs hinted so it should decode to an Object
[@map, ['a', 1], ['b', 2], ['c', 3]] # array of pairs hinted so it should decode to a Map
The nice thing about this approach is that JSON libraries could do as much or as little work in parsing out the annotations as they want: they could recognize the annotations and construct the referred-to type; or they could just pass back the array with the annotation expressed as an Annotation value.Map and set values are important for many applications. The fact that JSON doesn't have a way of denoting a map or set value (or anything else, but that's another issue) is a problem: it means there's no understanding common to all JSON consumers about what syntax denotes a map or a set. Shortcuts are taken, consumers disagree on the details, and the linked CVE is the result. Being able to reliably convey "this is intended to be a map", or "this is intended to be a set" is crucial for secure and robust interoperability.
And there are two (sensible) ways to add key/value pairs to a map: either keep the first occurrence of key X, or the last occurrence of key X.
Some JSON libraries pick one way; others pick the other.
Getting the two to interoperate is, as we see, not easy.
It would have been better for JSON to forbid duplicate keys, or to specify a mandatory first-wins or last-wins policy. Then there'd be no room for error.
It's fundamentally not the case that any data serialization format without support for map and set values allows for remote code execution. This CVE was CouchDB's auth code not handling some edge cases in the JSON standard.
Your argument that edge cases like this lead to problems is definitely a good one though. I'm more inclined to say that a lot of these kinds of issues are due to people thinking JSON is a simple, straightforward format when it definitely isn't -- and that's due to mostly two things:
- It fits on a business card!
- You can serialize/deserialize in one line in most implementations
Serialization/Deserialization is something you should always pay careful attention to. Making it a one-liner and advertising it as such was pretty irresponsible.
https://github.com/cblp/yaml-sucks
If you ignore the flame-inducing title, it's just a table showing how different implementations parse YAML input in very different ways.
maps can be implemented with 1 array and 1 map, with the keys being the hash of the object. the hash function should probably be written in the Json itself for completeness.
embedded beinary blobs already work. lookup GLB for an example on how to embed binary blobs in Json.
as others said, json already supports arbitrary precision.
Those aren't embedded binary blobs, those are string representations of base64-encoded binary blobs. Unless they come out of the JSON decoder as a byte array, they're not "working" as part of JSON; they're another standard on top of JSON.
(Also, the GP commenter probably wants them to be transmitted with 0 encoding overhead. Which can't really work while JSON is still JSON. But you can always use an alternative format which handles a superset of JSON's types, and which is binary, e.g. BSON or CBOR.)
> maps can be implemented with 1 array and 1 map, with the keys being the hash of the object. the hash function should probably be written in the Json itself for completeness.
I think you're fundamentally misunderstanding the thrust of the GP poster's request, here? They don't want to serialize a map in a way that is cheap to deserialize—different languages have different hashtable implementations so there's no way it could really work. What the GP poster (and many other people) want, is just to have something that encodes similarly to existing "object maps" (the ones with curly braces), but with a slight syntactic alteration so that they come out of the decoder as Maps, rather than as Objects.
- NaN/inf floats
- Comments
Schemas are defined somewhere, even if it's a poor and bug-ridden definition in the code. Ditto for the point that JSON doesn't count.
I wrote an article, On Schemas, to try and elaborate on this topic a bit.
This, and also I think a lot of people who preach the flexibility of schemaless anything really just wants a more flexible schema definition language, with a lot of "maybe" options and similar stuff.
> "So this program saves its data as 'S-expressions'." > "Oh, cool. Which dialect?" > "Uh, I don't know. It doesn't say in the README. It just says 'S-expressions'."
The answer can be CL, Scheme (lots of variations and dialects even here), the never-finished SPKI Sexps, OCaml sexps, something that the new dev on the team cooked up last Tuesday that vaguely resembles what they learned as "lisp" in college, or something else entirely.
On the other hand, if one were to say "this program uses R4RS S-expressions" (or, presumably, "CL S-expressions", but I haven't read the relevant bits of CLtL), you'd immediately be in a much nicer place than JSON can offer. Not only would you have a well-specified syntax for a reasonably broad range of data types, you'd have a useful equational theory as well. [ETA: Unless you want unicode. Doh. R6RS, maybe.] Ah, the impossible dream.
For most use cases, this isn't relevant, because the application expects a document of schema X and will provide said schema as well.