Gron – Make JSON Greppable
github.com
github.com
Why shouldn't I just use jq?
jq is awesome, and a lot more powerful than gron, but with that power comes complexity. gron aims to make it easier to use the tools you already know, like grep and sed.
gron's primary purpose is to make it easy to find the path to a value in a deeply nested JSON blob when you don't already know the structure; much of jq's power is unlocked only once you know that structure.
It's just operating at the wrong abstraction level, whereas gron is orders of magnitude easier to understand and _explore_.
I don't agree that jq's query language is obtuse. It's a DSL for JSON document trees, and it's largely unfamiliar, but so is xpath or any other DOM transformation language.
The same thing is said about regex.
My take is that "it's obtuse" just translates to "I'm not familiar with it and I never bothered to get acquainted with it".
One thing that we can agree though is that jq's docs are awful at providing a decent tutorial for new users to ramp up.
I also remember what when I first tried learning regex it was also very difficult. That is until I learned about finite state machines and regular languages, after that CS fundamentals class I was able to make sense of regex in a way that stuck.
Is there a comparable theory for jq's mini-language?
jq is basically a templating language, like Jsonnet or Jinja2. What jq calls a "filter" can also be called a template for the output format.
Like any template, a jq filter will have the same structure as the desired output, but may also include dynamically calculated (interpolated) data, which can be a selection from the input data.
So, at a high level, write your filter to look like your output, with hardcoded data. Then, replace the "dynamic" parts of the output data with selectors over the input.
Don't worry about any of the other features (e.g. conditionals, variables) until you need them to write your "selectors."
YMMV, but that's what's worked for me
in.json {"a":1,"b":2}
jq -c '{a}' in.json
{"a":1}
The . is the current stream, so if I just do ". , .", it's kind of pushing two streams along:
jq -c '.,. | {a}' in.json
{"a":1}
{"a":1}
Then, of course, say:
jq -c '{a, b, c: .}' in.json
{"a":1,"b":2,"c":{"a":1,"b":2}}
It was going through the . stream, and I pulled the . stream right back in while doing so.
So it kind of helps to keep straight in my head when I've kind of got multiple streams going, vs multiple values.
Someone (almost anyone) can probably explain better with formal theory, but I just kind of got a feel for it and kind of describe it like this.
jq has never come naturally. Every time I try to intuit how to do something, my intuition fails. This is despite having read its man page a dozen times or more, and consulted it even more frequently than that.
I've spent 20+ years on the Unix command line. I know my way around most of it. I can use sed and awk and perl to great effect. But I just can't seem to get jq to stick.
Aside, but there's a lot of times when "I know jq can do this, but I forget exactly how, let me find it in the man page" and then... I find jq's man page as difficult as jq itself when trying to use it as a reference.
Anyway, $0.02.
Edited to add: as a basic query language, I find it easy to use. It's when I'm dealing with json that embeds literal json strings that need to be parsed as json a second time, or when I'm trying to manipulate one or more fields in some way before outputting that I struggle. So it's when I'm trying to compose filters and functions inside jq that I find it hard to use.
I love jq, and without detracting from it, gron looks like an extremely useful, "less difficult" complement to it.
Just as an example, this just took me about a minute to get the data I wanted, whereas I probably spent a half an hour on it yesterday with jq:
curl -s https://static01.nyt.com/elections-assets/2020/data/api/2020-11-03/national-map-page/national/president.json | gron | grep -E 'races.*(leader_margin_votes|leader_margin_name_display|state_name)' | grep -vE 'townships|counties' | gron -ungron- a fellow read-only system
Perl shines for this use case (assuming it is present in the machines you are working with). It is slower than grep/sed/awk for most cases, but it is more powerful and better portable across platforms.
>converting YAML to JSON on the command line
check out https://github.com/bronze1man/yaml2json
Agreed.
For better or worse, when performance is not a concern in my scripts, I just shell out to "perl -pe" rather than trying to deal with grep, sed or awk.
It just works.
[1] https://blog.kellybrazil.com/2020/03/25/jello-the-jq-alterna... [2] https://github.com/kellyjonbrazil/jello
Here is a demo of a small script I wrote that shows jq results as you type using FZF: https://asciinema.org/a/349330 (link to script is in the description)
It also includes the ability to easily "bookmark" expressions and return to them so you don't have to worry about losing up an expression that's almost working to experiment with another one.
As a jq novice, I've personally found it to be super useful.
The slides are linked in the video description and at [2]. You'll need them because unfortunately the video is produced in such a way that the speaker video window often obscures important parts of his presentation.
[1] https://www.youtube.com/watch?v=_ZTibHotSew
[2] https://www.slideshare.net/mobile/btiernay/jq-json-like-a-bo...
Practically, most things you'd do with gron and grep, sed, awk, ... you could do using only jq as well. Jq comes with massive cognitive overhead though and has a bunch of very unpleasant gotchas (like silently corrupting numbers with abs > 2^53, although very recent jq graciously no longer does that iff you do no processing on the number).
I find jq pretty useful, but I have no love for it.
It really depends on what you want to do, and thus what think gron does.
If all you want to do is search for properties with a given value then yes, jq does that very well.
Unlike gron, jq even allows users to output search results as valid json docs. Hell, jq allows users to transform entire JSON docs.
However, if all you want to do is expand the JSON path at each symbol then I don't know if jq supports that usecase. But then again, why would anyone want to do that?
I like that jq's query expression syntax is command line (bash) friendly. My hunch is that xpath expressions would be awkward to work with.
I've done too much xpath, xquery, xslt, css selectors. For my own work (dog fooding), I settled on mostly using very simple globbing expressions. Then use the host language's 'foreach' equiv for iterating result sets.
Globbing's double asterisk wildcard is the feature I most miss in other query engines. https://en.wikipedia.org/wiki/Glob_%28programming%29
Looping back to command line xpath: there's always some impedance match between the query and host languages. IIRC, one of the shells, like chubot's oilshell or fish?, has more rational expression evaluation (compared to bash).
You especially see this with regexs. It's a major language design fail that others haven't adopted Perl's first class regex intrinsics. C# has LINQ, sure. But that's more xquery than xpath. And I've never liked xquery.
In other words, "blue collar" programming languages should have intrinsic path expressions. Whatever the syntax.
YMMV.
For example, against the public AWS IP address JSON document, it produces an output like
$ curl -s 'https://ip-ranges.amazonaws.com/ip-ranges.json' | jq -r '[path(..)|map(if type=="number" then "[]" else tostring end)|join(".")|split(".[]")|join("[]")]|unique|map("."+.)|.[]'
.
.createDate
.ipv6_prefixes
.ipv6_prefixes[]
.ipv6_prefixes[].ipv6_prefix
.ipv6_prefixes[].network_border_group
.ipv6_prefixes[].region
.ipv6_prefixes[].service
.prefixes
.prefixes[]
.prefixes[].ip_prefix
.prefixes[].network_border_group
.prefixes[].region
.prefixes[].service
.syncToken
This plus some copy/paste has worked pretty well for me.That jq query looks like an unwise Perl one-liner.
$ jq-structure my-file.jsonBTW: I have a D3 front-end dashboard/console for the app (not admin) that makes this a little bit harder, but D3 is pretty organized (and well-documented), if you can figure out what you are trying to do with it.
$ echo "{\"a\": 13911860366432393}" | jq "."
{
"a": 13911860366432392
}
$ echo "{\"a\": 13911860366432393}" | gron | gron -u
{
"a": 13911860366432393
}
I can now happily uninstall `jq`. I've been burned by it way too many times.I think structured data is so common now, that you have to invest in learning tools for processing it. Personally, I invested the time once, and it saves me every single day. In the past, I would have a question like "which port is this Pod listening on", and write something like "kubectl get pod foo -o yaml | grep port -A 3". Usually you get your answer after manually reading through the false-positives. But with "jq", you can just drive directly to the correct answer: "kubectl get pod foo -o json | jq '.spec.containers[].ports'"
Maybe it's kind of obtuse, but it's worth your time, I promise.
But then I squint a little bit at the default gron (not ungron) output, and that's actually what I see.
It seems to me that for your example use case, gron is at least useful to first understand the json structure before making your jq request. And, for simple use cases like this one, enough to replace jq altogether.
For Kubernetes specifically, "kubectl explain pod", "kubectl explain pod.spec", etc. will help you find what you're looking for.
Well, or you just do
kubectl get pod pod -o json | gron | grep port
and you will get the answer to the original question + the path.
For me, it was hours and hours and hours and days and days wasted with jq, before I found gron.
Not looking back.
json[0].commit.author.name = "Tom Hudson";
Now I need to escape brackets and dots in regex. Genius!I have 5 line (!) jq script that produces this:
json_0_commit_author_name='Tom Hudson'
This is what I call grep-able. It's also eval-able.> What if there's json object with commit and json_commit?
Then I'll use jq to filter it appropriately or change delimiter. The point is ease of use for grep and shell.
#!/usr/bin/jq -rf
tostream | select(length == 2) | (
( [ .[0][] | tostring | gsub("[^\\w]"; "_") ] | join("_") )
+ "=" +
( .[1] | tostring | @sh )
)It's not an ideal output format for sure, but it does meet some criteria that I considered to be desirable.
Firstly: it's unambiguous. While your suggested format is easier to grep, it is also lossy as you mention. One of my goals with gron was to make the process reversible (i.e. with gron -u), which would not be possible with such a lossy format.
Secondly: it's valid JavaScript. Perhaps that's a minor thing, but it means that the statements are eval-able in either Node.js or in a browser. It's a fairly small thing, but it is something I've used on a few occasions. Using JavaScript syntax also means I didn't need to invent new rules for how things should be done, I could just follow a subset of existing rules.
FWIW, personally I'm usually using gron to help gain a better understanding of the structure of an object; often trying to find where a piece of known data exists, which means grepping for the value rather than the key/path - avoiding many of the problems you mention.
Thanks for your input :) I'd like to see your jq script to help me learn some more about jq!
I very much like the look of gron for the simpler stuff!
You could also check argv[0] for if you were called via the `ungron` name. Then it would be as simple as a symlink, which is very easy to add at install/packaging time.
(I know it's fairly broadly known, but this is the "multicall binary" pattern: https://flameeyes.blog/2009/10/19/multicall-binaries/)
(Has bug reports, has PRs, it's not 'done'.)
I hope to be able to face dealing with people's issues and PRs soon.
All the best.
This looks perfect. Does one thing and does it well. I will be adopting this :-)
def flatten_tree: [leaf_paths as $path | {"key":$path | join("."), "value": getpath($path)}] | from_entries;
e.g. curl -s http://consul.service.consul:8500/v1/catalog/service/brilliant-service | jq -r 'flatten_tree'
I haven't felt the need to devise the reverse transformation yet, but it works great for grepping blobs of unknown structure $ echo '{"user": {"name": "Sam", "age": 40}}' | npx json-mask "user/age"
{"user":{"age":40}}
or (from the first gron example; the results are identical) $ gron "https://api.github.com/repos/tomnomnom/gron/commits?per_page=1" | fgrep "commit.author" | gron --ungron
$ curl "https://api.github.com/repos/tomnomnom/gron/commits?per_page=1" | npx json-mask "commit/author"
If you've ever used Google APIs' `fields=` query param you already know how to use json-mask; it's super simple: a,b,c - comma-separated list will select multiple fields
a/b/c - path will select a field from its parent
a(b,c) - sub-selection will select many fields from a parent
a/*/c - the star * wildcard will select all items in a fieldI have installed gron on all my development machines.
Will probably use it heavily when working with awscli. I'm not conversant enough in the jq query language to not have to look things up when writing even somewhat complex scripts. And I don't want to learn awscli's custom query syntax. :)
Thought at first that it might be possible to replicate gron's functionality by some magic composition of jq, xargs, and grep, but that was before I understood the full awesomeness of gron - piping through grep, sed maintains gron context so you can still ungron later.
Nice work, thank you!
2. Thesis: jq is cumbersome when used on a json input of serious size/complexity because upfront knowledge of the structure of the json is needed to formulate correct search queries. Gron supports that "uninformed search" use-case much better. Prove me wrong ;)
2. That's pretty much exactly why I wrote the tool :)
While that's only useful for picking out specific named keys without context, that's often good enough to get the job done. Added bonus is that json_pp and grep are usually installed by default so you don't have to install anything.
That said, gron certainly looks like it offers simplicity in cases where using jq would be too complicated.