FX: An interactive alternative to jq to process JSON
github.com
github.com
For example how would you take key k1 from a list of dicts [{k1: v1, k2: v2}, {k1: v3}]?
Do you mean something like:
.[].k1
Give it a try.jq does have a learning curve, but just like any query language, including SQL, first you need to learn the basics of the query language in order to get things to work.
In this case:
* you know that .[] iterates over objects, so you use it to unpack the root array,
* you know you get a stream of objects, thus from those you use the .k1 filter to get the values of each k1 key.
Here's jq's manual on basic filters: https://stedolan.github.io/jq/manual/#Basicfilters
After you get jq to filter out what you want, you can work on getting it to output results in whatever format you wish.
SQL is based on solid mathematical theory, relational algebra. I personally learned that (and tuple relational calculus) in college before learning SQL, which made it easier. It helps making it coherant. Is there something like this for jq? Often when people invent languages that are not based on solid theory, they tend to lack coherence. This can make learning them difficult if you're someone that relies on your mental model of how things "should" work, like I am.
It's a filter. You can name-drop math stuff and even mention monads and the like, but it's just predicates, maps, a reductions.
Also, I'm not aware of a single person who ever looked at relational algebra beyond the introductory lessons of a relational databases 101 course, and even then that stuff was mostly in the way.
I don't think my message was implying that jq is a worse (or better) tool for it. I was just explaining that for some people, tools with a theory behind are easier to learn and understand than tools without.
`gron` is great but doesn't seem to handle some (extreme-ish) situations that `jq` can, e.g. the json output from the fastnbt-tools. You either get a `token too long` error using `gron -s` because the input is too long (it's 90MB, that's fair) or you get only one set of outputs per key (iyswim) because they get overlapped in memory.
That sounds like a major bug. So it will silently skip data that you wanted?
It's definitely an oddness when you have multiple objects at the same level that aren't in an array but I guess the explanation there is "they should all be on their own individual lines as streaming json" which `gron` does handle correctly.
(echo '{"a":"23"}'; echo '{"a":"25"}') | gron -s
json = [];
json[0] = {};
json[0].a = "23";
json[1] = {};
json[1].a = "25";
> So it will silently skip data that you wanted?Yeah.
echo '{"a":"23"}{"a":"25"}' | gron
json = {};
json.a = "23";
The `-s` option doesn't help. echo '{"a":"23"}{"a":"25"}' | gron -s
json = [];
json[0] = {};
json[0].a = "23";Thanks for that tip!
My approach to the example would be to use `.[] | .k1` which I think does what you want, and like bash command line pipes, you can build up to it semi-interactively.
The bits I struggle with JQ often involve irregular json, where a value might be missing, or null, or a list, not sure what the idiomatic way to deal with that is if there is one.
And by the way, you can achieve live preview with any of these CLI tools by using fzf. This is the snippet for jql for example: `echo '' | fzf --print-query --preview-window wrap --preview 'cat test.json | jql {q}'` (substitute jql for jq or anything else)
P.S.: jql might seem dead, as there are no recent commits, but it's not. It's just finished.
For the special case you wrote as an example, where each object is just a single key-value, it's possible:
(object
"key" (pipe (keys) (0))
"value" (pipe ((keys)) (0)))That would be super, ta. `to_entries[]` is pretty much the major reason I've not managed to move off `jq` to anything else yet because it's just incredibly powerful in this situation.
The relevant jql snippet to solve this in the general case now is:
(pipe
(zip
(keys)
((keys)))
((keys)
(object
"key" (0)
"value" (1))))
It's not as terse as the jq equivalent - I'll probably add a way to create user-defined functions, so you can alias stuff like this to shorter forms - but that one will require more thought.Unfortunately my next issue is how do I iterate over an array of objects (like jq `.[]`)? I'm guessing it's maybe something to do with `range` but I don't know how many I have in order to fill in those indices and I can't do `(elem 0) ... (elem 1)` for the same reason.
Basically, you can think about the query as a composition of many functions which result in one big function taking in your JSON and outputting a new JSON.
When you do ("mykey") or (0) you dive in one level deeper. You can also transform what is that one level deeper by writing ("mykey" (mytransform)). There is a keys function which returns the list of keys or the list of indices, for the current object or list, respectively. And you can use those lists of indices for indexing purposes.
Thus, if you have an input list and want to transform it element by element, you can write ((keys) (my-single-element-transformer)). It gets the indices, uses them as an index, and transforms each object contained in the list.
So let's say you have a list of objects {"name": "abc", "surname": "xyz"} and would like to transform them into a list of {"abc": "xyz"}. You can write ((keys) (object ("name") ("surname"))). This goes over all elements and for each returns a single object with a key that is the name (it's actually a transformer/continuation which gets the name from the current object that we pass there) and value that is the surname.
You can also see that in the original "entries" query. It first zips the keys with the values, so for a list of {"mykey": "myvalue}, it will give you a list of lists ["mykey", "myvalue"]. Then it pipes that into another transform, which for each such pair creates an object {"key": "<first element of pair>", "value": "<second element of pair>"}.
The overall system isn't that straightforward at first, but playing around with it for a while should make it click and then it's easy to write even more complex queries.
'.[].k1' ╰─$ echo '[{"k1": "v1", "k2": "v2"}, {"k1": "v3"}]' | jq -r '.[] | .k1'
v1
v3
https://codefaster.substack.com/p/mastering-jq-part-1-59c1. parse a json value from stdin and set it as the initial result
2. for each function, apply the function to the result, and set the output as the result for the next function.
3. The final result is pretty printed on stdout.
Dunno if you'll see this given how many replies you already got, but rather than just dumping "how do you do that" here's a realization I had a while ago that made it way easier to understand:
jq's language is a series of filters/transformers more akin to bash pipes on a stream of data than anything else.
For example, just "." selects out the current object (and is needed to match the "root" at the start of the query), and jq pretty-prints the results (when to a terminal):
$ echo '[{"k1": "v1", "k2": "v2"}, {"k1": "v3"}]'
[{"k1": "v1", "k2": "v2"}, {"k1": "v3"}]
$ echo '[{"k1": "v1", "k2": "v2"}, {"k1": "v3"}]' | jq '.'
[
{
"k1": "v1",
"k2": "v2"
},
{
"k1": "v3"
}
]
There's only 1 matching element here, the outermost array. We want to go one deeper, so use "[]" to unwrap/flatten it: $ echo '[{"k1": "v1", "k2": "v2"}, {"k1": "v3"}]' | jq '.[]'
{
"k1": "v1",
"k2": "v2"
}
{
"k1": "v3"
}
jq is now iterating over 2 objects, so the next filter is the one where you select out the key you want. This can be done in two different ways for this example (per sibling replies): $ echo '[{"k1": "v1", "k2": "v2"}, {"k1": "v3"}]' | jq '.[].k1'
"v1"
"v3"
$ echo '[{"k1": "v1", "k2": "v2"}, {"k1": "v3"}]' | jq '.[] | .k1'
"v1"
"v3"
Note how I broke these up: The atoms are ".", "[]", and ".k1" - ".[]" isn't one of them, despite what it may look like at first glance when compared to ".k1". Some additional examples to show how these combine:The "unwrap/flatten" [] can be used multiple times when nested arrays are involved, with or without the pipe syntax, but only works on arrays. It errors if given something else:
$ echo '[[1,2,3],[4,[5,6]]]' | jq '.[]'
[
1,
2,
3
]
[
4,
[
5,
6
]
]
$ echo '[[1,2,3],[4,[5,6]]]' | jq '.[][]'
1
2
3
4
[
5,
6
]
$ echo '[[1,2,3],[4,[5,6]]]' | jq '.[][][]'
jq: error (at <stdin>:1): Cannot iterate over number (1)
$ echo '[[1,2,3],[4,[5,6]]]' | jq '.[] | .[]'
1
2
3
4
[
5,
6
]
$ echo '[[1,2,3],[4,[5,6]]]' | jq '.[] | .[] | .[]'
jq: error (at <stdin>:1): Cannot iterate over number (1)
Also notice how the "." is needed after the pipes; these are separate filters/transformations being chained together, so as a new rule it needs the same "." as with the first one.This one has the advantage of being natively understood by aws-cli, meaning you can pass a JMESPath to an AWS call and only receive the filtered / transformed result back.
I also created Jellex[1], which is a TUI built on Jello to assist with building the python queries.
Jello gives you the power of python but without all of the boilerplate, so it’s nicer to use in Bash scripts.
See https://github.com/simeji/jid/issues/66#issuecomment-4436718...
- FX "expand/collapse" functionality seems way better for exploring APIs whose shape you don't know
- jid is maybe marginally better for APIs where you have instant recall of the exact shape and need to rapidly query it
Overall, I like FX better because it provides feedback on your query faster.
I am grateful to the author(s) for creating it and I'll be using it instead of JQ whenever I need to wrangle APIs from the CLI.
[1] https://medium.com/@antonmedv/discover-how-to-use-fx-effecti...
shuf -n 1000 file
This is part of coreutils.
There's also jiq, which is a clone of jid (mentioned elsewhere) but with jq syntax
jq —-compact-output '.' | head -10 | foo
That is also useful for grepping to filter on records of interest.jq also has --stream for handling large inputs.
But then I discovered LINQPad[0] and, "The Legendary Dump".
Is how I usually do so, if you don't want the interactive mode drop the last pipe
(Its interactivity is keyboard- rather than mouse-based.)
[0] https://github.com/tomnomnom/gron
"Make JSON greppable!"
"gron transforms JSON into discrete assignments to make it easier to grep for what you want and see the absolute 'path' to it."
https://github.com/multiprocessio/datastation/tree/main/runn...
y2j () {
ruby -r json -r yaml -e 'puts JSON.dump(YAML.load(STDIN))'
}
Makes it easy to use json tooling for yaml, although of course it flattens out anchors etc. fx “code to eval” fx .comments[].authors[].names
You might want to take a look at this post[1] to see everything fx can do.[1] https://medium.com/@antonmedv/discover-how-to-use-fx-effecti...
If you are, then why comment?