Gron: A command line tool that makes JSON greppable
github.com
github.com
I maintain that any HTTP client functionality that supports enough options to be useful is complex and if it supports so few options that it's a toy, why include it?
There is non-negligible overhead in keeping track of what shell tools can make network requests, and not being correct and up to date in that area has been the cause of numerous bugs and security issues in tool sets and programs that include and utilize a component like this without realizing it.
To many, this might sound like some overly nit-picky complaint, but I maintain if you ask just about anyone that's been in the trenches as a sysadmin for more than a couple years whether they think saving the few characters it takes to use curl and pipe is a good trade-off for the possible unintended consequences of some developer shelling out to this without proper validation from some webapp, they'll tell you no.
This tool seems awesome, but keep it simple. It's not like it's a GUI tool and it's hard to pass data between programs. A pipe is perfect here.
I think you raise a good point; and I don't think it overly nit-picky.
Personally I don't find myself using the built-in HTTP client at all (and pipe the output of curl instead); but I know people who do. I ummed and ahhed about keeping this functionality for a while, and surveyed the—at the time fairly small—user base to figure out what I should do. What I found was a subset of users whose use-case was very different to my own; they were often running Windows, with no install of curl (which I believe is actually a default on Windows now?), and only wanted very basic functionality.
I'm definitely a proponent of the Unix philosophy, but I try my best to be pragmatic, especially where it can lower the barrier to entry for users.
- Only build that functionality on windows. Optionally, provide two packages for windows, one with URL fetching and one without it.
- Provide a flag that enables/disables the feature. Either default toit off and allow it to be enabled with a flag (and a helpful error message if oyu attempt to use it without a flag) or vice-versa.
- Do nothing. In the end, it's an entirely valid decision to leave it as it is unchanged. I thought it was something worth discussing, both for this tool and for the larger trend in general, but that doesn't mean the problem is so large that it makes the tool unusable until addressed.
In any case, thanks for making this. The simplicity and obvious usefulness of this tool means that as someone who deals with JSON quite often, generally in an archived form where I care about one or two entries in lists of hundreds, I imagine I'll be finding lots of use for it in the future.
Alas, no. It's just an alias to invoke-webrequest, or what ever their equivalent power shell incantation is.
The problem is that PowerShell 5.x still maintains the "curl" Invoke-WebRequest alias, and that captures `curl` on the PS command line before curl.exe out of the path. However this (and a bunch of other *nix-conflicting aliases) are removed with PowerShell Core, and annoying aliases can be deleted out of the Alias:\ PS-drive on PS 5.x and older.
Thanks!
By the way, do the Windows builds of gron automatically recognize and support UTF-16? If anyone does try piping curl output in PowerShell, that's what gron would receive -- PS automatically converts stdout and stderr output to UTF-16 because of its use of .NET String for stream I/O.
jq is awesome, and a lot more powerful than gron, but with that power comes complexity. gron aims to make it easier to use the tools you already know, like grep and sed. gron's primary purpose is to make it easy to find the path to a value in a deeply nested JSON blob when you don't already know the structure; much of jq's power is unlocked only once you know that structure.
$ jq -c tostream <<<'{"a":[{"b":2}]}'
[["a",0,"b"],2]
[["a",0,"b"]]
[["a",0]]
[["a"]]
$
However, filtering that and then reconstructing JSON from that is... not possible at this time: $ jq -c tostream <<<'{"a":[{"b":2}]}'|jq -crn 'fromstream(inputs)'
{"a":[{"b":2}]}
$
$ jq -c tostream <<<'{"a":[{"b":2}]}'|grep b|jq -crn 'fromstream(inputs)'
$
:(The reason is that tostream and fromstream can handle multiple top-level JSON texts, since jq normally does too, but then there's an ambiguity issue to resolve by having a sort of an object terminator. Filtering tostream's output with grep loses the terminators, and so fromstream cannot operate normally.
But it should be possible to define a function that does allow this, by, e.g., requiring just one top-level JSON text.
The other thing is that a path-based encoding that does not require quotes and commas would be handier -- tostream's output is itself JSON, so it's not shell-friendly. This is gron's brilliant innovation: it's got a path-based encoding of JSON that is easy to deal with in a shell script. (Mind you, I'm not sure that using brackets to denote array indices is all that easy to use, but the need to disambiguate object keys that look like numbers is critical. Also, there's an ambiguity as to keys that have embedded periods ('.') in them. And lastly, even gron can't shake off the string quotes for values.) That jq has the builtin functionality needed to do the same is not good enough if it doesn't actually do it out of the box.
And the way to get rid of ambiguity regarding object keys that contain periods or " = " (and also square brackets) is to escape them: ".." and " == " or similarly.
Example:
.foo.bar[0].baz == ..blah = this is a\ntwo-line string
where the last key in the path is "baz = .blah".Also, " = " is a bit annoying. I'd prefer ": ":
.foo.bar[0].baz: this is a\ntwo-line string
The the quoting rule for the special chars in keys can then be generic: double them. .foo.bar[0].baz[[5]]:: ..blah: this is a\ntwo-line string
Here the last key in the path is "baz[5]: .blah". Mind you, this is still not trivial to deal with in a shell script, so perhaps we need some other escaping mechanism -- one that doesn't reuse the escaped characters, such as \u escaping.But maybe I'm missing the benefit you're seeing, and it's not about searching?
Mind you, I'm sticking to jq, as I know it really well. But I'm thinking of other users here. I think the value of a path-based transformation of JSON is ease of use, which motivates me to think about making it even easier to use, such as by removing those quotes.
No need for sed for the path, use cut for that as well. cut -d'=' -f2- will remove the path (but leave a space).
In the end, you can accomplish it with the following, whichI think is fairly easy:
echo 'json[0].foo.bar.baz = "some string";' | cut -d'=' -f2 | sed -e 's/^\s*"//' -e 's/";$//'
For me, it's a toss up whether I would use that or Perl, since chances are I'm doing it as a first step in some other process, and I can just continue on in Perl for the rest of the process anyway. echo 'json[0].foo.bar.baz = "some string";' | perl -pE 's/^.*?"//; s/";$//;'
I find keeping the output as valid JS extremely useful though, since I can just paste a grepped entry into a developer console to get a valid object to play with on a page. That's cutting out a pipe to a js prettifier, pipe to less, search for identifying text, and careful cut and paste to get the enclosing block of text for what ends up being a semi-common action for me. On the other hand, I can get raw strings, but barely ever have need of that, and could fairly easily make an alias for that if it became common. .foo.bar[0].baz: this is a\ntwo-line string
The output of gron is perfect valid JavaScript. It wouldn't be that way with `:` as key-value delimiter.Mind you, one should not eval code to parse data. So I count this as a minus.
[0] https://www.npmjs.com/package/jsonsmash
[1] https://blog.tedivm.com/open-source/2017/05/introducing-json...
I've written a similar tool in Python, both for JSON and XML. Especially the JSON version was dead simple, probably fits on a single screen and took 15 minutes to test and write. Surely it didn't have any "features" but it does the job of letting me grep json.
Gron is probably 10x more versatile and actually comes with useful features but I'd really have to have pressing needs to do transformations of JSON on a regular basis to switch over.
The same applies to libraries in programming languages. There is a very vague threshold, depending on the expressiveness of the language and operating environment as well as the hardness of the problem itself, where it either makes sense to write your own library or reuse an existing one.
For some compare and contrast between: jsonpath, jq, and jp these were for getting AWS ec2 instance IDs:
cat foo.json | jsonpath -p $.InstanceProfiles.[*].RoleId
cat foo.json | jq .InstanceProfiles[].Roles[].RoleId
cat foo.json | jp InstanceProfiles[].InstanceProfileId
apologies to mobile users!In the end I wound up just using rq to convert the toml to JSON so I could use jq on it.
It was a while ago so I suppose it's possible I just used it too early in its life or something though.
I've found flattening JSON in this way not only useful for line based tools like grep, but also for understanding unfamiliar JSON. Sometimes it's nice to be able to see the whole path down to the value you're looking at.
cat xyz | python -mjson.tool | grep foo(json_pp is another pretty-printer that is likely to be already on your system.)
json.Host = "headers.jsontest.com";
json["User-Agent"] = "curl/7.43.0";
Why not use simple notation for eveything without '.' in key?The output is designed to be valid JavaScript, which doesn't allow certain characters in unquoted object keys, like the dash in User-Agent.
Using JavaScript's rules for quoting keys makes it a lot easier to specify the grammar (and therefore write the parser); and makes it trivial to 'parse' the output using JavaScript should you want to.
There must be _some_ rules in place for when to quote the key (e.g. when there is a dot, equals sign, square brace etc in the key name), so I see no reason to adopt something custom and potentially error-prone when a known-good set of rules already exists.
Hope that answers your question well enough!