An Introduction to JQ
earthly.dev
earthly.dev
curl -s https://api.github.com/repos/stedolan/jq/issues?per_page=2 | jq '[ .[] | { title: .title, number: .number } ]'
curl -s https://api.github.com/repos/stedolan/jq/issues?per_page=2 | jq '[ .[] | { title, number } ]' map({title, number})
instead [ .[] | <something> ]For example, map() operates on a stream of arrays. That kind of doubling of iteration can be confusing at first.
Thanks.
I always look for "jq for csv" but perhaps I should just convert to json, then back to csv with jq
The author probably has internalized more of the manual than they realize, and maybe improved at least one explanation; FTA:
> map(...) let’s you unwrap an array, apply a filter and then rewrap the results back into an array. You can think of it as a shorthand for [ .[] | ... ] and it comes up quite a bit in my experience, so it’s worth it committing to memory.
From https://stedolan.github.io/jq/manual/#map(x),map_values(x) :
> map(x) is equivalent to [.[] | x]. In fact, this is how it's defined. Similarly, map_values(x) is defined as .[] |= x.
Note here the casual introduction of the update assignment operator, '|='
`.[]` takes a list and turns it into a sequence consisting of each element of that list.
`| x` applies the filter `x` to that sequence, turning it into a new sequence.
The outermost `[ ]` builds a list from that new sequence.
But I'd really like to see a discussion about the tool that the host website promotes: Earthy. It a build system, which is a family of tools that I've always hated, but it seems to be pretty decent. Is anybody using it?
I'm off to find HN threads on Earthly.
One convenient tip I discovered after reading this article and trying out the command is that
jq 'map({ title: .title, number: .number, labels: .labels | length }) | map(select(.labels > 0))'
can be refactored into jq 'map({ title: .title, number: .number, labels: .labels | length } | select(.labels > 0))'
or in other words, map(filter1) | map(filter2) == map(filter1 | filter2).With a relatively large jq program like that, it is critical that the main recursive loop run efficiently, so it's annoying that there's no way to detect whether tail call optimization was applied, other than benchmarking. It would also be nice if object values were lazily evaluated so that it would be possible to create ad hoc switches.
Compare that with „gron“ ( https://github.com/Deitar13/gron ), which is arguably not as powerful and clunky, but it allows me to compose and tie into the other unix tools way better.
It‘s mainly for that reason I use it more than jq these days for ad-hoc analysis.
Part of the blame is definitely on the fact that the "idiomatic" way to filter columns in a shell pipeline is to invoke AWK: `awk '{print $3,$5}'`. Similarly for JQ. Virtually every sysadmin and programmer gets introduced to these languages as Unixy tools when in fact they are antithetical to the Unix philosophy.
The result is ending up with overcomplicated "production" pipelines (curl|sed|awk|jq) when you really could be writing one far more coherent, maintainable, scalable C, Go, Python, etc. program with their standard libraries.
I realized it after wasting multiple hours debugging problems at work in scripts that used jq or AWK or similar tools and most of the time the problem was solved by quoting randomly things, except when discovering that there is the edge case that I didn't considered and the program broke, again.
Now when I have that sort of problems I don't even bother trying to fix them, I just rewrite the whole script in python (they are usually small scripts so it's a question of 15 minutes most of the time). And writing new script in bash is banned (except particular cases).
Also there is the concept of portability, most people assume that everyone has a way to install that tools because they have on their system, it's not that simple, and while putting things in production or on the CI it breaks because jq is missing. And good luck with Windows, by the way.
One nit: ‘jq -r’ seems like more fundamental than a sidenote, especially considering it’s a cli tool. That could just be how I use it though (as glue between json and bash).
Another good tool in that family is ‘jtbl’. It provides table output, which is useful for cut, awk, sed and column.
* https://github.com/tomnomnom/gron - make json greppable
* https://sr.ht/~gpanders/ijq/ - interactive jq
fzf --print-query --preview-window wrap --no-clear --preview 'cat file.json | jq {q}'
you can also pipe a curl in the above, but that will mean a lot of (slow) requests, so I have this snippet saved for running fzf --preview with jq on something from a web service curl -L https://datahub.io/core/covid-19/r/worldwide-aggregate.json > /tmp/foo && echo '' | fzf --print-query --preview-window wrap --no-clear --preview 'cat /tmp/foo | jq {q}'
now if you write something like .[0]["Confirmed"]
you will get a live preview of the result as you typeI did not know about `jiq`, so thanks for that tip, it looks like it does the same or something similar, but without the extra cruft of storing a temp file
The author of the tool has also written a guide [2] and recorded a screencast [3] about the tool.
[1] https://github.com/antonmedv/fx
[2] https://medium.com/@antonmedv/discover-how-to-use-fx-effecti...
However powerful jq may be, the comment above, which has exactly been my experience with jq, summarizes quite nicely the biggest hurdle with this tool.
99% of my jq use looks like this:
cat file.json | jq . | <regular list of unix filters>
and the remaining 1% is straight cut and paste from google / stackoverflow that may or may not end up doing what I want.jq's DSL is inscrutable
This is the second time someone recommended this on HN, so this time, I did go and have a look.
Really nice indeed, thanks for the tip.
From the project's description: Miller is like awk, sed, cut, join, and sort for name-indexed data such as CSV, TSV, and tabular JSON
Here's a simple way to list and browse github issues of a given user/repo:
#!/usr/bin/env sh
browse_url() {
firefox http://github.com/$1/issues/$2
}
issue=$(curl https://api.github.com/repos/$1/issues |
jq -r 'map([(.number|tostring), .title] | join(" | ")) | join("\n")' |
dmenu -i -l 10 |
awk "{print \$1}")
browse_url $1 $issue
Although the `| join("\n")` part could be done in a more idomatic way with just `[]`, sometimes the manual way are still clearer to me: map([(.number|tostring), .title] | join(" | "))[] [(.number|tostring), .title] | join(" | ")
I would write: "\(.number) | \(.title)"
which IMO is more readable in cases where you have specific values you want to put in specific places, as opposed to a list of unknown length which you want joined (eg. I would still use join("\n") in your example).I didn't know that one could build an arbitrary string like that inside a map.
Thanks a lot for that, I agree it looks better!
Most useful jq cli flag is -f. - take the jq script from file
Most useful tutorial for learning to manipulated your OpenAPI spec https://apihandyman.io/api-toolbox-jq-and-openapi-part-1-usi...
I have a few utilities involving JQ that I wrote.
For structured logs, I have jlog. Pipe JSON structured logs into it, and it pretty-prints the logs. For example, time zones are converted to your local time, if you choose; or you can make the timestamps relative to each other, or now. It includes jq so that you can select relevant log lines, delete spammy fields, join fields together, etc. Basically, every time you run it, you get the logs YOU want to look at. https://github.com/jrockway/json-logs. Not to oversell it, but this is one of the few pieces of software I've written that passes the toothbrush test -- I use it twice a day, every day. All the documentation is in --help; I should really paste that into the Github readme.
I am also a big fan of using JQ on Kubernetes objects. I know what I'm looking for, and it's often not in the default table view that kubectl prints. I integrated JQ into a kubectl extension, to save you "-o json | jq" and having to pick apart the v1.List that kubectl marshals objects into. https://github.com/jrockway/kubectl-jq. That one actually has documentation, but there is a fatal flaw -- it doesn't integrate with kubectl tab completion (limitation of k8s.io/cli-runtime), so it's not too good unless you already have a target in mind, or you're targeting everything of a particular resource type. This afternoon I wanted to see the image tag of every pod that wasn't terminated (some old Job runs exist in the namespace), and that's easy to do with JQ: `kubectl jq pods 'select(.status.containerStatuses[].state.terminated == null) | .spec.containers[].image'`. I have no idea how you'd do such a thing without JQ, probably just `kubectl describe pods | grep something` and do the filtering in your head. (The recipes in the kubectl-jq documentation are pretty useful. One time I had a Kubernetes secret that had a key set to a (base64-encoded) JSON file containing a base64-encoded piece of data I wanted. Easy to fix with jq; `.data.THING | @base64d | fromjson | .actualValue | @base64d`.
JQ is something I definitely can't live without. But I will admit to sometimes preprocessing the input with grep, `select(.key|test("regex"))` is awfully verbose compared to "grep regex" ;)
Stuff I do with it:
- prepare json request bodies for curl commands by constructing json objects using environment variables
- grab content from a deeply nested json structure for usage in a script
- extract csv from json
- pretty print json or ndjson output curl ... |jq -C '' | less -r. I actually have an alias set up for that.
The syntax is a bit hard to deal with. I find myself copy pasting from stack overflow a lot when I know it can do a particular thing but just can't figure out how to do it.
Just pipe curl to node (https://github.com/jareware/howto/blob/master/Replacing%20jq...) or Python or Ruby or whatever you already know!
[0] https://gist.github.com/Checksum/72d927471c76c76c46418b3ee88...
Everytime I want to use it I end up searching up examples of the syntax and not quite getting it right.
``` cat file.json | jq 'keys' | grep ependencies ```
This will list the keys, sometimes is really helpful with a big json that you don't know the schema, but you have the intuition that some key should be there.
jq 'keys | map(select(test("ependencies")))' file.json
map(select()) is a pretty useful construct to loop over stuff and pick out the interesting parts. cat node_modules/some_lib/package.json | jq '.version'[1] https://mosermichael.github.io/jq-illustrated/dir/content.ht... [2] https://news.ycombinator.com/item?id=22626080
curl -G 'https://api.example.com' \
-d foo=bar \
-d baz=whee
Is the same as curl 'https://api.example.com?foo=bar&baz=whee'
That can make all the quoting hell a bit easier, as you're doing it one param at a time.??
It can really be overwhelming when you realize they control all of the major institutions in our country.
no need to memorize
> The frequency illusion is that once something has been noticed then every instance of that thing is noticed, leading to the belief it has a high frequency of occurrence
This is clearly a flavor-of-the-month tool that has been getting much coverage on these sorts of sites lately, why impune the guys mental function?