Show HN: Catj – A new way to display JSON files
github.com
github.com
jq -j '
[
[
paths(scalars)
| map(
if type == "number"
then "[" + tostring + "]"
else "." + .
end
) | join("")
],
[
.. | select(scalars) | @json
]
]
| transpose
| map(join(" = ") + "\n")
| join("")
'
EDIT: Got the string quoting and escaping.EDIT 2: For those who want to save this script, you can put just the jq code in an executable file with the shebang:
#!/usr/bin/jq -jf jq -r '
tostream
| select(length > 1)
| (
.[0] | map(
if type == "number"
then "[" + tostring + "]"
else "." + .
end
) | join("")
) + " = " + (.[1] | @json)
'
EDIT: For those who want to save this script, you can put just the jq code in an executable file with the shebang: #!/usr/bin/jq -rf jq -r '
( tostream
| select(length > 1)
| (
.[0] | map(
if type == "number"
then "[" + tostring + "]"
else "." + .
end
) | join("")
)
+ " = "
+ (.[1] | @json)
+ " |"
),
"."
' ( jq "$(sed 's/$/ |/;$a.')" <<< '{}' )
As in: catj example.json \
| ( jq "$(sed 's/$/ |/;$a.')" <<< '{}' ) \
> original.jsonBTW the input json can be null, so -n works (also using process substitution):
jq -nf <(sed 's/$/ |/;$a.') jq -c --stream '
. as $in
| select(length == 2)
| (
$in[0] | map(
if type == "number"
then "[" + tostring + "]"
else "." + .
end
) | add
) + " = " + ($in[1] | tostring)'
Using `--stream` allows jq to start before parsing the entire json file. In my experience, a 700mb json file can take up 5gb of ram in either jq or python -m json. .movie.name = "Interstellar"
.movie.year = 2014
.movie.is_released = true
.movie.else = "Christopher Nolan"
.movie.cast[0] = "Matthew McConaughey"
.movie.cast[1] = "Anne Hathaway"
.movie.cast[2] = "Jessica Chastain"
.movie.cast[3] = "Bill Irwin"
.movie.cast[4] = "Ellen \\\\ Burstyn"
.movie.cast[5] = "Michael Caine"
You're outputting: ".movie.name = Interstellar"
".movie.year = 2014"
".movie.is_released = true"
".movie.else = Christopher Nolan"
".movie.cast[0] = Matthew McConaughey"
".movie.cast[1] = Anne Hathaway"
".movie.cast[2] = Jessica Chastain"
".movie.cast[3] = Bill Irwin"
".movie.cast[4] = Ellen \\\\ Burstyn"
".movie.cast[5] = Michael Caine"
Another point is how the strings at the right of the `=` are displayed. They should be quoted. The reason why they're not is because you piped the second element to `tostring` instead of `@json`.A better version of your suggestion would've been:
jq -r --stream '
select(length > 1)
| (
.[0] | map(
if type == "number"
then "[" + tostring + "]"
else "." + .
end
) | add
) + " = " + (.[1] | @json)
'
The use of `length > 1` instead of `length == 2` is a minor point, but if a future version jq decides to sometimes put 3 elements in these arrays, your filter would ignore those when we're likely to also want those. `length > 1` ensures what we need, that there are at least the elements that we're going to be using, while `length == 2` might filter some of those out, even if it's not right now.Your use of `add` is neat, though. I wouldn't have thought of that.
https://gist.github.com/fernandoacorreia/4b67a41bbe227654868...
Maybe not, but I'm pretty sure every system supports a single arg. And very few (none?) support more.
From https://pubs.opengroup.org/onlinepubs/9699919799/utilities/V...
> If the first line of a file of shell commands starts with the characters "#!", the results are unspecified.
There is no more specification for "#!/usr/bash" than "#!/usr/bin/jq -jf"
The exec page provides even fewer words about how to interpret shebangs if you thought perhaps I was linking to the wrong portion of the posix spec
jq -r '
tostream
| select(length > 1)
| (.[0] | map("[" + @json + "]") | join(""))
+ " = " + (.[1] | @json)
'
And this other option with the blacklist patterns: jq -r '
tostream
| select(length > 1)
| (
.[0] | map(
if type == "number" or (tostring | test("[@-]|^[0-9]|^else$"))
then "[" + @json + "]"
else "." + .
end
) | join("")
) + " = " + (.[1] | @json)
'
(The blacklist here is non-exhaustive, but an example.) jq -r '
tostream
| select(length > 1)
| (
.[0] | map(
if tostring | (
test("^[A-Za-z$_][0-9A-Za-z$_]*$")
and (
. as $property
| ["if", "else"] | all(. != $property)
)
)
then "." + .
else "[" + @json + "]"
end
) | join("")
) + " = " + (.[1] | @json)
'
You whitelist against what the syntax allows for identifiers and then you blacklist reserved keywords. Writing it this way makes it easier to verify for correctness when comparing with the ECMAScript Specs. This is still a non-exhaustive blacklist and the whitelist regex lacks allowed unicode characters. jq: error: syntax error, unexpected INVALID_CHARACTER, expecting $end (Unix shell quoting issues?) at <top-level>, line 3:
jq -j '
jq: 1 compile error
This, uh, doesn't work for me on jq-1.5.1.tostream | select(length > 1) | ( .[0] | map( if type == "number" then "[" + tostring + "]" else "." + . end ) | join("") ) + " = " + (.[1] | @json)
`#!/usr/bin/jq -rf ` with tostream wrapper in code works fine
jq "--stream -rf" path/to/script
and jq doesn't know of any one option called "--stream -rf".I haven't seen the discussions around these design decisions in the different OSes, but I imagine the crux of the matter is that you have to pick somewhere to stop, and where you chose to stop is largely arbitrary.
I mean, you can have the OS interpret shebangs with multiple arguments, but then you'll want to be able to put spaces in these arguments, so you'll want quoting, and then you'll want to put special characters like newlines inside, so you'll want escaping, etc.
The OS can implement all these things in execve()'s logic, but it might also be preferable to keep the logic simple in the interest of avoiding security-harming bugs. You know, less code, less bugs, less vulnerabilities.
If --stream had a single letter option equivalent, you could stick it together with the other ones. However, since it doesn't, your only option to make a portable script is to use a shell shebang like #!/bin/bash, and then do:
exec jq --stream -rf ...
You might feel that this single argument restriction sucks and is definitely inferior to any implementation of multiple argument shebangs. I don't know if macOS shebangs support quoting, but if they don't and simply split on spaces, then I can tell you they can't do hacky stuff like writing code in a shebang like this:> https://unix.stackexchange.com/questions/365436/choose-inter...
Granted, it's bad practice, but a little cool nevertheless.
$ docker inspect 620f55df9177| structure.sh |grep -i addr
.[].NetworkSettings.GlobalIPv6Address
.[].NetworkSettings.IPAddress
.[].NetworkSettings.LinkLocalIPv6Address
.[].NetworkSettings.MacAddress
.[].NetworkSettings.Networks.bridge.GlobalIPv6Address
.[].NetworkSettings.Networks.bridge.IPAddress
.[].NetworkSettings.Networks.bridge.MacAddress
$ docker inspect 620f55df9177| jq .[].NetworkSettings.IPAddress
"192.168.0.2"The project linked to is from 2014 with last update in 2015.and it is on NPM...
What is left to say? Thank you!
The idea of flattening, grepping, then reverting sounds very appealing and sounds like a better fit for me.
I don't think you really need neither `.` nor `-C`. Just `jq` seems to do the same colored output of the input by default.
`-C` would be required when piping because most of the time (with the exception of piping into less) when stdout is not a terminal, it doesn't make sense to include terminal color escape sequences. You'd end up with those codes in your files, and grep would be looking at them for matches, for example.
`.` would be required when passing the file as an argument instead of stdin, because jq interprets the first argument as jq-code. If you don't include `.` it would interpret the filename as jq-code.
I do honestly think jq is a cool and powerful tool. I also appreciate little things like auto-color when appropriate--git also does this. Git also uses your pager, which might trivialize my personal use case.
There are cases when you have some complicated json and just want to search for stuff. Then you use grep + gron.
There are cases when you want a complete json processing tool. Then you use jq.
You can probably simulate each approach with the other approach, but the code needed to this is just too tedious to write. So you use whatever tool fits your use case.
It still feels like there must be something in between, some way to make queries with json more naturally, than with jq, yet with enough power.
It might help to recognize how it's influenced by shell languages and XPath, if you're familiar with those.
https://github.com/twpayne/flatjson
The flat format is great for diffs:
--- testdata/a.json
+++ testdata/b.json
@@ -1,5 +1,6 @@
root = {};
root.menu = {};
+root.menu.disabled = true;
root.menu.id = "file";
root.menu.popup = {};
root.menu.popup.menuitem = [];
@@ -9,8 +10,5 @@
root.menu.popup.menuitem[1] = {};
root.menu.popup.menuitem[1].onclick = "OpenDoc()";
root.menu.popup.menuitem[1].value = "Open";
-root.menu.popup.menuitem[2] = {};
-root.menu.popup.menuitem[2].onclick = "CloseDoc()";
-root.menu.popup.menuitem[2].value = "Close";
-root.menu.value = "File";
+root.menu.value = "File menu";(I haven’t run it, but a skim of the code suggests that this tool will turn `{"foo.bar": "baz", "foo": {"bar": "baz"}}` into `["foo.bar"] = "baz"` and `.foo.bar = "baz"`, resolving the separator ambiguity in a pretty JavaScripty way.)
The text streams that are processed line-by-line by dozens or hundreds of line-based tools are immensely powerful and universal. It's all Unix heritage and often overlooked by fancy modern designs that more often follow a fashion rather than root themselves in substance.
Surely text streams have their share of limitations like everything else but in practise you can retrofit nearly anything into line-based text streams and get an immediate productivity multiplier by being able to apply a whole array of established tools to process that data. Proof of that power is that it has been worthwhile to write converters to and from text and other formats. Not only you can find translators to turn various hierarchical or object-oriented formats into text but you can even convert a PNG into text and back (with SNG).
Text streams are like roads with lanes. They're ages old, they're pretty good at separating and guiding traffic, and they're somehow suboptimal in several senses yet rarely can anyone point out a single, clear practical improvement on laned roads, not to mention a system for containing traffic flows that is superior to them.
$ augtool -r . -L --transform 'JSON.lns incl /catj-eg.json' <<< 'print /files/catj-eg.json'
/files/catj-eg.json
/files/catj-eg.json/dict
/files/catj-eg.json/dict/entry = "movie"
/files/catj-eg.json/dict/entry/dict
/files/catj-eg.json/dict/entry/dict/entry[1] = "name"
/files/catj-eg.json/dict/entry/dict/entry[1]/string = "Interstellar"
/files/catj-eg.json/dict/entry/dict/entry[2] = "year"
/files/catj-eg.json/dict/entry/dict/entry[2]/number = "2014"
/files/catj-eg.json/dict/entry/dict/entry[3] = "is_released"
/files/catj-eg.json/dict/entry/dict/entry[3]/const = "true"
/files/catj-eg.json/dict/entry/dict/entry[4] = "director"
/files/catj-eg.json/dict/entry/dict/entry[4]/string = "Christopher Nolan"
/files/catj-eg.json/dict/entry/dict/entry[5] = "cast"
/files/catj-eg.json/dict/entry/dict/entry[5]/array
/files/catj-eg.json/dict/entry/dict/entry[5]/array/string[1] = "Matthew McConaughey"
/files/catj-eg.json/dict/entry/dict/entry[5]/array/string[2] = "Anne Hathaway"
/files/catj-eg.json/dict/entry/dict/entry[5]/array/string[3] = "Jessica Chastain"
/files/catj-eg.json/dict/entry/dict/entry[5]/array/string[4] = "Bill Irwin"
/files/catj-eg.json/dict/entry/dict/entry[5]/array/string[5] = "Ellen Burstyn"
/files/catj-eg.json/dict/entry/dict/entry[5]/array/string[6] = "Michael Caine"
[0] $ ls .../share/augeas/lenses/dist/|wc
221 221 2867https://blog.tedivm.com/open-source/2017/05/introducing-json...
I've used JSONExplorer for this purpose, but it is web based and doesn't handle files this large.
Extending the filesystem metaphor to JSON data and re-using the same commands strikes me as a great idea.
Did another project inspire you, or did you come up with the concept yourself?
Have you done a Show HN yet?
As you've already implemented most relevant commands (cd, ls, cat), it would probably be easy to make a FUSE version using fs-fuse / fuse-bindings
function cason(x){
switch(x[0]){
case "movie": switch(x[1]) {
case "name" : return "Interstellar";
case "year" : return 2014;
case "is_released": return true;
case "director" : return "Christopher Nolan";
case "cast": switch(x[2]){
case 0: return "Matthew McConaughey";
case 1: return "Anne Hathaway";
case 2: return "Jessica Chastain";
case 3: return "Bill Irwin";
case 4: return "Ellen Burstyn";
case 5: return "Michael Caine";
}
}
}
} def license(kernel):
return {"Linux": "GPL",
"FreeBSD": "BSD",
"NT": "Proprietary"}[kernel]https://sqlite.org/json1.html#jtree
SELECT big.rowid, fullkey, value
FROM big, json_tree(big.json)
WHERE json_tree.type NOT IN ('object','array');Related, if you want more of a csv-style, see JSONLines. aka "newline-delimited JSON"
And JSON is (almost) a perfect subset of yaml.
I've been using csv lately. It's reputation is overstated.
What I like is that it's far more compact than yaml or json and trivially pulled into sqlite for ad-hoc queries.
Or spreadsheets, to work out a plan, and then MySQL.
Another is that it makes streaming processing a little easier. Once you have a line, you know you can attempt to process it, and you can shard processing on newlines without a full YAML processor. Tools that work on newlines or tools that can just split on lines can handle the first level of JSON Lines output.
deno install catj https://deno.land/std/examples/catjson.ts --allow-read
https://github.com/zacharyvoase/jsonpipe
It includes 'jsonunpipe'.
So you could grep part of the JSON and still get a JSON back.
``` echo '{"a": 1, "b": 2}' |grep b| jsonpipe | jsonunpipe
#{""b": 2} ```
This may look cute but it is horrific when dealing with large configs and you have to reconstruct all the structure in your head.
Also, when you have a format that nests using brackets, braces and parenthesis, you can get help from the editor. This format does not give you that.
I'm not a huge fan of JSON (and the above mentioned format was invented because none of us were fans of XML at the time), but it turns out that both XML and JSON are actually easier to work with in practice than this format. Not least because there is ample tooling for JSON (and XML).
The lesson I learnt: I may hate XML (or in this case JSON), but finding an alternative that is better is not easy.
jq can also do the same thing, with more flexibility. And it is possible to combine with bash alias to make it indistinguishable from catj
... or you know, you could put the jq script in an executable file and add a shebang like
#!/usr/bin/jq -jf
or #!/usr/bin/jq -rf
In my opinion, aliases should mostly be used to add default options only. Not really to insert whole scripts into them.That said, if you're stuck dealing with bad JSON like this with low signal to noise this is a decent way to redisplay it.
var json = JSON.parse('{"my": "json"}'); (function printRecursively(ob, _keys = []){ _.map(ob, (val, key) => { var k = _.isNumber(key) ? '[' + key + ']' : '.' + key; var keys = _keys.concat(k); if (_.isObject(val)) printRecursively(val, keys); else console.log(keys.join('') + ' = ' + typeof val + ' ' + val); }); })(json);
EDIT: how do I markup code on HN?
Indent each line with four spaces. Please please keep line length very short (under forty?) as HN's pre tags are absolutely not mobile friendly and can trash the entire page.
Edit: actually it seems to at least scroll within the comment div on overflow now, that's a huge improvement!
Here's a Python implementation of gron https://github.com/venthur/python-gron
And for purely aesthetic, nitpicky reasons I think the leading period in each line is redundant.
I could also see it useful for just finding the index of a particular array entity.
This is a nice CLI based alternative
Here's an example from a couple of days ago:
https://twitter.com/mikemaccana/status/1141706132823695362?s...
XML is native to nothing. It's annoying to use everywhere.
UNIX tools just don't compare. Sure, they stick around, because sometimes they're the easiest tool for some job, until they aren't and you regret starting out with them.
In a few years, Javascript and Javascript-compatible languages are likely to be bigger than ever before. A whole generation of developers has been trained mainly on them. Billions of dollars have been invested into the ecosystem. Whether we like it or not, that's the reality. For that reason alone, JSON will stick around.