Mastering Jq: Part 1
codefaster.substack.com
codefaster.substack.com
Yikes:
> wget -O - -q 'https://reddit.com/r/unixporn/new.json' | jq
I just want the links:
> wget -O - -q 'https://reddit.com/r/unixporn/new.json' | jq '..|.permalink? | select(.)'
And every time I try to learn, I get lost in a maze of a manpage:
https://manpages.debian.org/jq
I mean there are useful tips in there, but I rarely find what I need. I usually find it simpler to write a Python script (which ships with `json` because it's "batteries included") and operate on lists and dicts the normal way...
It's longer, but at least I don't need to learn a new programming language (if we can call jq that...)
The reason it's popular partly because it's available from both Centos/Debian based. Just use `yum` and `apt`. Most of time, you maybe limited with what have on a container. And `jq` may even pre-install by DevOps.
Second is it play well with unix tool, pipe and friend.
I think the only way would be a long web page with a massive list of examples to "solve problem X".
But the DSL is ...byzantine, to say the least.
jq has looping, branching, recursion (with tail recursion optimization), variable assignment, modules, I/O... Yeah, it's a programming language.
I've made jql[0] for that reason.
Check it out, it's less featureful, but much simpler with a uniform lispy syntax.
I don't think this is uncommon. People generally reach for jq for simple path expression evaluation against JSON objects, and never get deeper into it than that. It seems like the kind of thing you could re-implement yourself in an afternoon. However, as soon as you start taking advantage of some of the more complex functionality--say, a program like `.entries[] | select(.size > 1024) | .name`--there's a disquieting feeling of "what the hell is actually happening here?"
I took some time a while back to really dig into how jq programs work, and was surprised to discover how deep and powerful jq really is, while being built out of some very simple fundamental building blocks. Most of the jq builtin functions are actually implemented in jq itself[1]; even more of them could be, but aren't for the sake of performance. And the power and self-consistency of jq made sense when I found out that the creator Stephen Dolan is an "actual" Computer Scientist (in the academic sense) with extensive programming language research experience[2].
The main obstacle I see for most people is the lack of of an accessible introduction to how jq actually works when it comes to streams, function calls, generators, and backtracking. They're explained somewhat in the docs and on the project wiki[3][4][5], but in a fairly blunt way that assumes existing familiarity with terminology and concepts. It takes some effort to learn initially, but once you understand it, everything falls into place. I'm hopeful that subsequent installments in this "mastering" series can explain how the jq model works in an approachable way.
[1] https://github.com/stedolan/jq/blob/master/src/builtin.jq
[2] http://stedolan.net/research/
[3] https://stedolan.github.io/jq/manual/#Advancedfeatures
[4] https://github.com/stedolan/jq/wiki/Advanced-Topics
[5] https://github.com/stedolan/jq/wiki/Internals:-backtracking
May I suggest "gron"? As compared to jq it is simpler (good!) has less features (good!) has no DSL (good!) and it is less powerful (good!, in many cases). It is a tool that "expands" json into standalone lines, and you can further process them using the standard tools grep, sed, cut, sort, awk...
EDIT: for example, if you have this json:
{
"outdir" : "out3",
"data" : [
{"img" : "img_01.jpg"},
{"img" : "img_02.jpg"}
],
"roi" : {
"x" : 150,
"y" : 150,
"w" : 700,
"h" : 700
},
"margin_h": 20,
"margin_v": 5,
"tile_size" : 300,
"resolution": 0.5
}
running it through gron produces this outdir = "out3"
data[0].img = "img_01.jpg"
data[1].img = "img_02.jpg"
roi.x = 150
roi.y = 150
roi.w = 700
roi.h = 700
margin_h = 20
margin_v = 5
tile_size = 300
resolution = 0.5
that some people find really convenient to deal with.This is not to shame the developers but to voice our issues with it since this seems to be the prevailing sentiment, perhaps some brave soul will try and simplify it for human brains.
Yq [1] syntax for that is much better, and yq can accept JSON files too.
Any more alternates to Python or Node that are lightweight to add to a project? I'd rather use Node but that's a massive dependency for a non-Node project.
In python, you can have a jq like experience with
def f(data):
# put code here
data = json.load(sys.stdin)
output = f(data)
print(json.dumps(output))I think the examples for any() are any(true, false) and any(true, true). Fair play but not super helpful if I’m new and trying to understand how this fits into a typical jq filter.
I would like to see a document that talks about jq from this more abstract perspective and what capabilities exist in this context. For example, can I select two arbitrary branches? Can I merge / move branches? Once we have a grasp of tree manipulation capabilities that are made available through the DSL we can lookup the syntax. Can I select a node and apply filters to it's children? There are so many ways in which a tree can be manipulated and I personally need a higher level description of what I can accomplish with jq.
Spending one hour or so actually learning jq is a good time investment.
echo '' | fzf --print-query --preview 'cat example.json | jq {q}'
Here's a rough example of it in action: https://asciinema.org/a/y4WGyqcz1wWdiyxDofdPXKtdC echo '{"k1": [{"k2": [9]}]}' | jq '.k1 | .[0] | .k2 | .[0]'
is equivalent to echo '{"k1": [{"k2": [9]}]}' | jq '.k1[0].k2[0]'
I kept waiting for the author to explain this, but they did notWhen teaching though, I'd rather make the it more visually explicit that jq commands are a sequence of filters. With the former notation, a | explicitly tells the reader that it's the end of one filter and the start of a new one, the latter is more implicit.
From a teaching perspective it is probably best to start with the long-form.
However, there are quite a few typos in the examples... even the very first one is missing a quote and won't parse if cut-and-pasted as is.
I do my scripting mostly in Go now and it's much easier to structure my code and grow it over time. I sometimes use Python, but Go is more flexible for what I do. I can compile an exe and drop it onto a host and run it, without worrying about VM and library versioning (or in Bash's case, making sure jq is installed and all the Unix utils have compatible versions).
I don't know what you mean by "memory-management-free" though, Go is still a garbage collected language.
Also I recently learned Go's CLI parser is pretty weak compared to argparse.
(bash doesn't care about errors either, by the way.)
And handling errors in concurrent code is consistent--still passed as values. Languages with implicit stack-unwinding exceptions have inconsistent ways to deal with the errors because, in concurrent programs, error handling doesn't end once the stack is unwound, because there are lots of concurrent stacks. You then usually catch the exception near the top of the stack and pass it to another stack...as a value.
So yes, there is some boilerplate involved in that. But as a codebase grows, I appreciate that the language encourages me and others to be intentional about how errors are handled, and also provides a consistent mechanism for handling them in concurrent code (concurrency is a big reason I use Go to begin with).
After getting super frustrated with the documentation while trying to accomplish stuff that should have been straight forward, I created jsling so I could pipe output through node for JavaScript one-liners. Anything moderately complicated that doesn't need to be portable; I just use that.
EDIT: I should mention that by "moderately complicated" I mean stuff that starts to get into the realm of joins, correlated sub queries, and the like.
https://programminghistorian.org/en/lessons/json-and-jq#the-...
My approach uses very simple ideas and is heavily based on JS internally. This is simple once you understand it [3] (thus saner DSL) and gives the whole power of JS in your hands.
I've also prepared a basic comparison jq vs jsqry based on some examples in article [4].
It's worth noting that currently the CLI tool [1] is written in Java using Graal VM polyglot native image compilation. Thus HUGE executable size of 95 MB (sic!) because it bundles JS engine. I'm considering rewriting this to QuickJS by Fabrice Bellard [5]. This should make it MUCH smaller.
[1] https://github.com/jsqry/jsqry-cli
[2] https://github.com/jsqry/jsqry
[3] https://jsqry.github.io/#filtering
[4] https://gist.github.com/xonixx/d6066e83ec0773df248141440b18e...
I love how nested lists inside objects can expand into a particular one with the .[] operator.
For example:
{ a: [{b: 5}, {b: 3}], c: 5} can be transformed into: [{ b: 5, c: 5}, {b: 3, c: 5}] using jq '{c: .c, b: .a[].b}'
For heavily nested XMLs I can get a nice flat output.
In a future post, I'll cover how to use jq not just for json and xml but any data format.
Now, admittedly anywhere we have Node available, it pretty much negates a lot of it. But even with node, it's sometimes easier to just throw in a couple jq calls to extract data.
I've written some relatively complex programs with it, e.g. this one [1] I just wrote to manipulate the output from neuron [2] to generate index files based on tags. It's probably overkill/not as efficient as it could be, but I wanted to be able to expand upon it later.
I think once you grok the basics, which can be a bit confusing at first, you can accomplish some amazing things with the help of a few of the more advanced features like reduce, to_entries/from_entries, etc.
I really wish it was more actively developed, I feel like it actually has potential to be a semi-general-purpose functional programming language.
But I find jq more pleasant than most for lots of JSON-to-JSON transformations, and Powershell worse than not only jq but also Node or Python for that task.
Yes, and that’s great. I don't prefer jq for downloading JSON (which it doesn't do), but, as I said, for JSON-to-JSON transformations.)
`tail -f my-log-file.json | jq`
This can be great if you've setup Nginx to log as JSON.