The Pretty JSON Revolution
ohler.com
ohler.com
`python3 -m json.tool somefile.json` or `cat foo.json | python3 -m json.tool` will print it in "one line per node" format. 3.9 introduces a --sort-keys switch for sorted objects also.
> An object is an unordered collection of zero or more name/value pairs, where a name is a string and a value is a string, number, boolean, null, object, or array.
That is, you could rely on this and but be aware of it, is my point. Firefox, for example, will happily take an object with duplicates and report only the last one.
That said, I agree that being aware of this is important if you're emitting JSON. You'd think nobody would ever address a JSON object by its ordinal position, but programmers are lazy and worse, think they're clever. :)
As of Python 3.6 (in theory not guaranteed until Python 3.7), key order is preserved when reading and writing. That's a consequence of the fact Python dictionaries can now remember the order of insertion.
It's true that it doesn't support duplicate keys though (unless you pass in a different class to the object_pairs_hook parameter of loads() to replace its use of dict).
"The JSON syntax does not impose any restrictions on the strings used as names, does not require that name strings be unique, and does not assign any significance to the ordering of name/value pairs. These are all semantic considerations that may be defined by JSON processors or in specifications defining specific uses of JSON for data interchange."
https://www.ecma-international.org/wp-content/uploads/ECMA-4...
(I have actually encountered a order-dependent JSON-subset parser before, but to my mind, that code is broken)
JavaScript itself will also sort object keys if they are numeric, so `{a: "a", c: "c", b: "b", "1": 1};` will be transformed to `{1: 1, a: "a", c: "c", b: "b"}`.
The main apps I've seen that depend on JSON structure are for hashing, which would also be broken by whitespace / linebreak variances in pretty-printers.
In my experience jq already does pretty the output. Maybe I'm missing something in your comment.
I hoped to find jq as a gem/module/library but I was disappointed. After days of searching and trying different things, I honestly could not find any powerful library or API for traversing and searching JSON.
Something like:
open file.json | select colors | each { ^echo $it.hex }
is much nicer than jqjson_pp < somefile.json
The only thing I don't like is that it doesn't process commandline arguments. You have to pipe the file in. It is also fairly strict, I've run into a number of malformed JSON files that it rejects but other parsers would accept. Naked TRUE/FALSE statements are one thing it hates that are super common, especially from places like Google.
This seems to be a bad idea. The JSON language spec has ORDERED object members. But the order is arbitrary (precisely the one given in the JSON string) and does not have to be the lexicographic.
Sorting the object members by default would introduce problems whenever the order matters to the consumer of the JSON.
It's unclear: At one point it says "An object is an unordered set of name/value pairs." while in the actual grammar it is ordered:
object
'{' ws '}'
'{' members '}'
members
member
member ',' membershttps://www.ecma-international.org/wp-content/uploads/ECMA-4...
> An object is an unordered collection of zero or more name/value pairs, where a name is a string and a value is a string, number, boolean, null, object, or array.
"whenever the order matters to the consumer of the JSON" should be never.
More pragmatically, regardless of what the spec says, a ton of JSON tooling assumes the order doesn't matter and relying on it would be a big mistake.
That's the great thing about specs -- if you don't like what one says, there's always another to support your position. :)
(This kind of text is why I prefer the ECMA document -- it's clearly written to be a normative standard, rather than as a field-report and Request for Comments.)
JSON is a serialization format. Its components inherently have a serial order. You can't change this any more than you can legislate the value of pi to be 3.
Consider just how how many data formats are ultimately defined as a "serial stream of characters" -- and then consider how few of those you would practically use for marshalling a general data structure.
Also, there's the problem of different JSON libraries behaving differently. Such as using unordered hashmaps as an internal data representation for parsed content, making compliance difficult.
Don't mess with the order yourself, but don't assume other tooling will respect it.
Personally, I will adopt the first version of JSON that lets me insert a flipping comment!
Agreed. But this does not mean that a tool should break it.
My assumption would be that
fn(parse(pretty_print(someJSONString)))
should always evaluate to the same as fn(someJSONString)
(for all functions fn)JSON object literals expressly follow a different rule (are expliclty unordered) per the IETF specs and have no specific significance to order at the JSON level though some might conceivably be introduced in ancillary specifications or tooling per the bigECMA spec.
JSON is syntactically a subset of JS but not semantically identical. Unambiguous order would require an array of one-entry objects in JSON.
> An object structure is represented as a pair of curly bracket tokens surrounding zero or more name/value pairs. A name is a string. A single colon token follows each name, separating the name from the value. A single comma token separates a value from a following name. The JSON syntax does not impose any restrictions on the strings used as names, does not require that name strings be unique, and does not assign any significance to the ordering of name/value pairs. These are all semantic considerations that may be defined by JSON processors or in specifications defining specific uses of JSON for data interchange.
https://www.ecma-international.org/wp-content/uploads/ECMA-4...
This wording allows a particular implementation to define its own meaning to the order of key-value pairs, or even to produce a multimap.
{
"foo": "this is a comment about foo",
"foo": "actual value of foo that overwrites the comment"
}
The trick is that the second value value of foo overwrites the first. But, clearly, sorting would would wreak havoc here (if the value was used in the sort key). ;)A fun fact about MongoDB is it will actually store that JSON, both duplicate keys. The implication is that whatever MongoDB client you're using, that maps Mongo data to dictionaries/maps, is not capable of representing all valid MongoDB documents. It's important to recognize that Mongo may be storing data your client will not be able to access.
I learned this when the Python client was showing one value for a key, and the Ruby client was showing another value for the same key, and neither client was showing the whole document.
False. “An object is an unordered collection of zero or more name/value pairs, where a name is a string and a value is a string, number, boolean, null, object, or array.” [emphasis added][0]
The normative text has the "real" answer, and the real answer is that it's basically undefined behavior. It starts by saying "The names within an object SHOULD be unique", and then elaborates:
An object whose names are all unique is interoperable in the sense that all
software implementations receiving that object will agree on the name-value
mappings. When the names within an object are not unique, the behavior of
software that receives such an object is unpredictable. Many implementations
report the last name/value pair only. Other implementations report an error
or fail to parse the object, and some implementations report all of the
name/value pairs, including duplicates.
JSON parsing libraries have been observed to differ as to whether or not they
make the ordering of object members visible to calling software.
Implementations whose behavior does not depend on member ordering will be
interoperable in the sense that they will not be affected by these
differences.
https://tools.ietf.org/html/rfc8259#section-4Unlike some RFCs, which clearly and explicitly delineate normative from informative material, RFC 8259 does not, but the text you cite is on its face informative rather than normative: it does not specify what an implementation MUST or SHOULD do, or what the object model IS, it describes the variety of preexisting implementations (based, correctly or not, on prior specifications) that are in the wild.
At this point I honestly take XML over JSON where I have a choice because of CDATA and comments.
{
"colors": [
{ "color": "black", "hex": "#000", "rgb": [ 0, 0, 0 ] },
{ "color": "red", "hex": "#f00", "rgb": [ 255, 0, 0 ] },
{ "color": "yellow", "hex": "#ff0", "rgb": [ 255, 255, 0 ] },
{ "color": "green", "hex": "#0f0", "rgb": [ 0, 255, 0 ] },
{ "color": "cyan", "hex": "#0ff", "rgb": [ 0, 255, 255 ] },
{ "color": "blue", "hex": "#00f", "rgb": [ 0, 0, 255 ] },
{ "color": "magenta", "hex": "#f0f", "rgb": [ 255, 0, 255 ] },
{ "color": "white", "hex": "#fff", "rgb": [ 255, 255, 255 ] }
]
}That's an incredibly XML-ified version of a color table. I can clearly see the tags now. Can't just do a look up of a color color, instead I would have to iterate over the members or store it in a different data structure.
Why even use JSON? Blech.
"colors": { "red":{"rgb":"fff"}", ... }
You key by color names, that seems obvious. But what if you want to look up a color by hex? Now you have to look through them all.
What if this list is actually an order list of colors for different headings? And they can repeat? Then indexing by index is exactly what you want.
Point is, you don't know the reason behind the data structure.
(That said, I do wish the json had a standardized way to remove redundancy for objects that always follow the same structure. One list of property names, and then everything in arrays.)
As you would with the original one as well.
> Then indexing by index is exactly what you want.
A much better use case. That said, the list order is kinda fragile, having an explicit row identifier might be worth adding. Especially since all of the columns are being explicitly called out instead of in their own list.
That said, given there are just shy of 17 million possible RBG combinations, and a small fraction those are of named colors, I'd personally continue to optimize for the named color case.
Surely you wouldn't think of sending a piece of JSON over the wire with both name:color and rgb:color, regardless of whether that is what the recipient wants to operate on. You just have to let it unmarshall the data into whatever form it needs.
That said, my comment may a bit misdirected as a result, in which case: Mea Culpa.
And before you argue that dictionaries can still be iterated in order, you better check the sibling threads where people are arguing you shouldn’t rely on that.
color | hex | rgb ║
red | #f00 | [3] ║
black | #000 | [3] ║
yellow | #ff0 | [3] ║
green | #0f0 | [3] ║
cyan | #0ff | [3] ║
blue | #00f | [3] ║
magenta | #f0f | [3] ║
white | #fff | [3] ║
I would be nice to inline the rgb column here.Numbers to the right make it much more pleasant to my eyes
{
"colors": [
{ "color": "black", "hex": "#000", "rgb": [ 0, 0, 0 ] },
{ "color": "red", "hex": "#f00", "rgb": [ 255, 0, 0 ] },
{ "color": "yellow", "hex": "#ff0", "rgb": [ 255, 255, 0 ] },
{ "color": "green", "hex": "#0f0", "rgb": [ 0, 255, 0 ] },
{ "color": "cyan", "hex": "#0ff", "rgb": [ 0, 255, 255 ] },
{ "color": "blue", "hex": "#00f", "rgb": [ 0, 0, 255 ] },
{ "color": "magenta", "hex": "#f0f", "rgb": [ 255, 0, 255 ] },
{ "color": "white", "hex": "#fff", "rgb": [ 255, 255, 255 ] }
]
}You mean:
{
"colors": [
{ "color": "black" , "hex": "#000", "rgb": [ 0, 0, 0 ] },
{ "color": "red" , "hex": "#f00", "rgb": [ 255, 0, 0 ] },
{ "color": "yellow" , "hex": "#ff0", "rgb": [ 255, 255, 0 ] },
{ "color": "green" , "hex": "#0f0", "rgb": [ 0, 255, 0 ] },
{ "color": "cyan" , "hex": "#0ff", "rgb": [ 0, 255, 255 ] },
{ "color": "blue" , "hex": "#00f", "rgb": [ 0, 0, 255 ] },
{ "color": "magenta", "hex": "#f0f", "rgb": [ 255, 0, 255 ] },
{ "color": "white" , "hex": "#fff", "rgb": [ 255, 255, 255 ] }
]
} {
"colors": [
{ "color": "black" , "hex": "#000", "rgb": [ 0, 0, 0 ] },
{ "color": "red" , "hex": "#f00", "rgb": [ 255, 0, 0 ] },
{ "color": "yellow" , "hex": "#ff0", "rgb": [ 255, 255, 0 ] },
{ "color": "green" , "hex": "#0f0", "rgb": [ 0, 255, 0 ] },
{ "color": "cyan" , "hex": "#0ff", "rgb": [ 0, 255, 255 ] },
{ "color": "blue" , "hex": "#00f", "rgb": [ 0, 0, 255 ] },
{ "color": "magenta", "hex": "#f0f", "rgb": [ 255, 0, 255 ] },
{ "color": "white" , "hex": "#fff", "rgb": [ 255, 255, 255 ] },
]
}More importantly, what’s more readable to you?
That said, it is just another pun on the text as art thing. In that it doesn't really scale, and you are going to upset someone by not having a codified tool for automatically doing this. (I don't recall seeing align-regex in any popular tool.)
[ { foo: a bar: 123.45 }
{ foo: abc bar: 6.7 } ] foo.bar.baz = 10
.biz = 12 // foo.bar.biz
..boz.baz = 31 // foo.boz.baz
etc. It basically combines really brittle context-sensitive grammar production with complete lack of greppability.I mean, at least before one has proven a tool's ubiquitous use, use a longer name.
jq just got lucky but I don't think it was because of its name ;).
Treating the colons as white space, as you've done with the commas, will move you one step closer to The Correct Answer™.
SEN is new. After dealing with broken JSON due to commas missing or one at the end of an array and some of the team using Javascript this was a way of sucking in the broken JSON and fixing it.
Postel tried to warn us.
> ...nice having the extra reminder that the left side of the colon is a key and the right a value.
Totally. IMHO: whitespace, formatting, delimiters are for humans. The parsers can do without. With some exceptions, like your examples of quoting strings to remove ambiguity.
Remark and Unified are some well-known projects that wooorm maintains.
JSON.stringify(JSON.parse(require('fs').readfileSync('myfile.json')),null,2);
Valid JSON:
{
"key1": "hello",
"key2": "world"
}
You could "trick" a JS formatter to format it by wrapping with a fake function, etc. Some minimum valid JS: json({
key1: "hello",
key2: "world",
});
MongoDB has some JS libraries that use similar tricks to use JS parsers for their shell query format (which is similar to JSON). For example, around line 597: https://unpkg.com/browse/ejson-shell-parser@1.1.1/dist/ejson...Is that a convention I'm not aware of? Seems a little obtuse and unnecessary, why not just accept two arguments? One less arbitrary usage detail to remember.
I like the idea that the incompatible format is off by default.
How does SEN deal with numbers-encoded as string? is it something like .4 ? that's a bit confusing
And the source for that homepage is here: https://github.com/treenotation/treenotation.org
Always open to PR!
The whole idea of pretty notation is automatically inserting non-significant whitespace to make it look nice. Step 2, "one line per node", inserts spaces and newlines. Step 4, "human style" strategically removes some of those so the lines look nice -- the 2nd level dict has lots of content, so it was split across multiple lines... while the 3rd level dict has fewer data, so it all fits on one line.
As opposed to this, Tree Notation is all about single canonical representation. So whitespace is significant, and you can never add or remove it to make output look nicer. You do whatever your schema tells you, and I hope you like many short lines.
So what they are really talking about is just pretty code. Their favorite examples utilize alignment (tree notation does that better—every tree doc is ismorphic to a spreadsheet and you don't have to align things to the left spine, and their are grid langs that don't do that).
The colors et al are called "secondary notations" and again Tree Notation can't be beat. Adding secondary notations is simple. Here's an example: https://www.youtube.com/watch?v=vn2aJA5ANUc
(I am talking about "Human Style with Colors" here -- this is the cool part, and I don't really expect SEN to take off except to display things on the terminal)
That tabular-like alignment was generated automatically -- I can take any existing JSON data source and the program will automatically make it look nice while not requiring any changes in the consumers or producers of the data.
Compare to Tree Notation, for example this code: https://jtree.treenotation.org/designer/#standard%20iris has this block with has a clear structure:
sepalLengthCell
extends floatCell
sepalWidthCell
extends floatCell
petalLengthCell
extends floatCell
petalWidthCell
extends floatCell
speciesCell
enum virginica versicolor setosa
highlightScope constant.language
This looks pretty ugly to me. There is clearly the table-like structure, but it is hard to see, because each line is split in 2. If this were JSON/SEN, I could make it look nicer: sepalLengthCell: { extends: floatCell }
sepalWidthCell: { extends: floatCell }
petalLengthCell: { extends: floatCell }
petalWidthCell: { extends: floatCell }
speciesCell: {
enum: [virginica, versicolor, setosa]
highlightScope: constant.language
}
See how it's all aligned now and how structure comes out? And all at zero effort from my part, it was all computer generated? But with Tree Notation, the above is invalid -- it has different meaning, so the compiler won't accept it. You have to use much uglier vertical method, with all the newlines.And Pretty JSON can also adapt to display width. Someone with large fonts or small display can request 80 characters wide output, and "speciesCell" will be wrapped. Someone with huge display can request output 250 characters wide, and "speciesCell" will be column-aligned with others. Another thing which is pretty impossible in Table Notation without a lot of work.
No. It takes JSON only. Which is one format out of 10,000 (though a popular one).
> There is clearly the table-like structure
It is a directed acyclic graph structure.
I won't disagree that perhaps in certain older tools oj may be better in certain situations. But Tree Notation (or more generally 2D/3D languages where positioning is the only thing used for syntax) are the future. The key thing to keep in mind is that without the colors in the last 2 examples OJ is not very useful or pretty. So to make oj nice you need to start adding parsers which are necessary for secondary notations. Then once you start adding secondary notations, 2D/3D langs make that orders of magnitude easier.
Is that iris grammar document in designer a good example? Because it does not really show why is it better than something like SEN or JSON. The text alignment is awkward, and it does not fit well on the screen. The colors are there, but I am not convinced that having keywords be a different colors outweights the ugly formatting. And I am not sure what "secondary notations" are, but I am guessing they are not present in the document?
The idea of the schemas are good. The world could use more editors with context-aware syntax highlighting and auto-completion. However, the data model you have chosen ("TreeNode=tuple(string, list[TreeNode]") doesn't map well to any programming language. And the text serialization you have chosen -- having a single canonical representation that is using whitespace as the only syntax element -- un-nesseary restricts the data that can be represented.
I think switching to a more conventional data model and text serialization will significantly increase uptake of your project, as well as make it more aesthetically pleasant. Because I seriously doubt anyone can call your existing grammar programs "nice looking".
You keep saying that "Tree Languages are the future"
but I don't see any impressive examples of it.
Binary notation as an idea was worked through for ~250
years before we had very impressive examples of it in
computers, so I'm pretty happy with the examples so far
given that I'm ~10 years in, and ~4 years since publication.
So just at a high level the short term game common in the
rest of tech isn't something I'm interested in playing.Beyond what's out there, I've seen the results from thousands of experiments in everything from assemblies to compiler compilers to declarative data notations, and from software to hardware and everything in between, so the amount of data I have dwarfs what everyone else has seen. While I totally get that it's not raining buckets now, and in fact people feel barely a drizzle, I have quite a dataset that there are big clouds on the horizon.
Is that iris grammar document in designer a good example?
It's mildly neat. Here's how we used an early version of
that a couple of years ago to publish synthesized data for a
GWAS EOPEC study
(https://github.com/breckuh/eopegwas/blob/master/mockData/cli...).
Tree Notation will become the standard way to describe data
schemas and make synthesis a breeze. outweights the ugly formatting
The formatting can be described as this: minimal. In fact,
the most minimal. If you think minimal is ugly, than we
probably won't come to an agreement. Keep in mind though
that you can write code to project Tree programs in whatever
way you want. I won't disagree with the statement "Tree
Notation doesn't work as well with my existing tools as
other langs", but if you go back to stuff from 2017 and look
at the trajectory, you'll see that Tree notation tooling has
improved remarkably and in a couple more years you'll see
stuff that just isn't possible with 1-D langs. doesn't map well to any programming language
Do you know Lisp? Tree Notation maps to S-Expressions
without parens. un-nesseary restricts the data that can be represented.
From the paper (2017): "Prediction 1: no structure will be
found that cannot serialize to TN." your existing grammar programs "nice looking".
To each their own. I think in the long run simplicity lasts. Also, the bigger idea isn't Tree Notation, but the idea of 2-D and 3-D languages https://longbets.org/793/Pretty JSON is inevitably for either logging or config files, and YAML is better at both of those.
1. First notice that there is a world of difference between what users want and what they are willing to achieve. Know this more than anything else. People will ask for all kinds of shit, and.... A wish list is not a fully explored business requirement with known sub-tasks and test cases. A simple ask can become something worthy of a different independent project.
2. Too subjective. Everybody has subtle different personal preferences. In some cases the inability to support some edge case of some language will cause certain users to have an emotional episode. WTF. This is free software providing a convenience that you can easily live without.
3. A lot of work. You have to be very clear about what language, grammar, class of languages, or other various of characters you are willing to support. For example there is HTML then there are about billion trillion different HTML template schemes each with their own syntax and inside that syntax is a wildly different language than the surrounding HTML.
4. Carve out a measurable portion of your life. This is an investment of time you will never get back. Writing a code beautifier is far more work than it sounds. First, you need a parser. If one does not exist for the language you wish to support in the language or format of your tool you will need to write one. Be careful though, because that parser will have to support conventions that are unique to beautification and not necessarily useful elsewhere. In the case of the HTML example above you will need multiple different parsers that can achieve a nesting of parse trees or achieve harmony of a uniform parse tree beloved by all languages. This is achievable, as I have done it, but good luck.
5. Maintenance. There are always new edge cases, new languages, new grammars, new features and your users will want them all. Set hard boundaries.
------
With the amount of work required you will begin to ask yourself some basic life questions:
Does this tool bring me more money or a better job? Does it bring me prestige AND satisfy a craving for attention? Does it improve my work, as in other real work outside your beautification tool?
In my case, for a while, the tool did allow me access to better jobs with increased pay. It demonstrated I could do things many other developers could not and that I was willing to dedicate some absurd about of effort into something people actually used. But, that will only take your career so far after which you are just spinning your wheels and burning time.
When I got further in my career I realized I wasn't beautifying my code ever. I had no need for the tool I was maintaining and despite continuous maintenance by me the tool started to decay, because the requirements had grown out of control and I was no longer an end user.
Your final example is just approaching a JSON -> YAML converter. If your complaint about your chosen human readable serialization format is that it isn't human readable enough, then switch to something more inherently human readable instead of writing tools to temporarily transform it.
YAML is “a superset of JSON”, yes, but there are two separate meanings to that:
• YAML has alternative syntactic sugar for expressing the same underlying JSON-equivalent semantics (sort of the same as Avro being canonically a binary compact expression of underlying JSON — in both cases, libraries for the codec expect JSON-encodable data structures as #encode input, and produce JSON-encodable data structures as #decode output)
• YAML has its own semantics (like node type annotations, or references) that JSON doesn’t have, such that documents that use these are no longer transposable into JSON.
I love bullet point #1. I hate bullet point #2.
Personally, I wish there was a name for the reduced subset of YAML that is still a “syntactic superset of JSON”, but which has none of the extended semantics of bullet-point #2.
Many systems that “consume YAML” already actually require their documents to be this “strictly-JSONifiable YAML”! Kubernetes, for example: it might seem to expose a YAML manifest API, but actually, internally, it does everything in JSON. All the resources in k8s etcd are stored in canonicalized JSON. The k8s controller just prettifies that JSON to YAML on its way out to you; and uglifies it back to JSON when you send it in. Which means that any YAML features that don’t survive that translation, can’t be used.
IMHO, if YAML hadn’t been designed with any extended semantics, but instead had strictly targeted being a “sugared alternative encoding of JSON”, I think everyone would have switched to sending YAML in place of JSON a long time ago. Browsers would have likely added YAML parsing as well.
But those added semantics are just so much extra work for everybody. Type annotations are source of so many vulnerabilities in programs that were unaware their input could “reach in and do things” through those types; and yet many YAML parser libs don’t have any flag to restrict them from decoding these type annotations (i.e. no way to “defuse the bomb.”) References change the entire way you have to write a YAML parser, disallowing some types of parsing grammar altogether, meaning you might no longer have access to the first-class parsing solution of your language runtime; meaning that for many runtimes, the YAML codec lib for that runtime is much slower — and memory-intensive! — than the JSON codec lib for the same runtime. Etc.
Honestly, if we could all agree on a name for “strict, JSONifiable YAML”, and create libraries that only parse/validate/accept that subset of YAML while rejecting the higher-level semantics, those libs—and that interchange format—would be immediately more popular than YAML. The time for this to happen hasn’t passed! We still have a chance!
It's too ugly for humans (too many quotes, too many escape characters, and no comments) and too texty for machines.