jq 1.7
github.com
github.com
I love JQ so much we implemented a subset of JQ in Clojure so that our users could use it to munge/filter data in our product (JVM and browser based Kafka tooling). One of the most fun coding pieces I've done, though I am a bit odd and I love writing grammars (big shoutout to Instaparse![1]).
I learned through my implementation that JQ is a LISP-2[2] which surprised me as it didn't feel obvious from the grammar.
[1] https://github.com/Engelberg/instaparse
[2] https://github.com/jqlang/jq/wiki/jq-Language-Description#:~....
.[].commit | select(.author == "Tom Hudson")
which basically says "find all commits by Tom Hudson" in the input.`.[]` iterates all the values in its input (whether the input be an array or an object). `.commit` gets the value of the "commit" key in the input object. You concatenate path expressions with `|`, and array/object index expressions you can just concatenate w/o `|`, so `.[]` and `.commit` can be `.[] | .commit` and also `.[].commit`. Calls to functions like `select()` whose bodies are path expressions are.. also path expressions.
Perhaps the most brilliant thing about jq is that you can assign to arbitrarily complex path expressions, so you can:
(.[].commit | select(.author == "Tom Hudson")) = "Anon"
The syntax is strange probably because of this trying to make path expressions so trivial and readable.jq programs get hard to read mainly when you go beyond path expressions, especially when you start doing reductions. The problem is that it resembles point free programming in Haskell, which is really not for everyone.
The other thing is that jq is very much a functional programming language, and that takes getting used to.
https://github.com/flox/flox/blob/019095f8bc40e49abc8e5cd0b1...
commits = [elt.commit for elt in data if elt.commit.author = "Tom Hudson"]
json.dump(commits, sys.stdout)
Definitely not as straightforward... would be nice to have a bit more affordances for path expressions in Python.
data = json.load(sys.stdin)
commits = [
e["commit"]
for e in data
if e["commit"]["author"] == "Tom Hudson"
]
json.dump(commits, sys.stdout)Has nothing to do with arrays, it has to do with the fact that Python dicts with string indexes and Python objects with properties are different things, unlike JS where member and index access are just different ways of accessing object properties.
> Or maybe I'm spoiled from using typed languages and cannot see the ingenuity of the python/javascript/other-untyped-hyped-lang api authors that it solves?
This isn't an untyped thing, this is a JavaScript (and thus JSON) and Python have type systems (even if they usually don't statically declare them) and those type systems and thus the syntax around objects are different between the two.
P.S. `some.get("A", {})["B"]` is bad programming habit because there might be a list on `some["A"]`
jello '[e.commit for e in _ if e.commit.author == "Tom Hudson"]'
Jello let’s you use python syntax with dot notation without the stdin/stdout/json.loads boilerplate. import jmespath
import json
doc = json.load(sys.stdin)
print(jmespath.search("[?commit.author == `Tom Hudson`].commit", doc))
I wish it had won over jq because JMESPath is a spec with multiple implementations and a test suite where jq is... well jq and languages have bindings not independent implementations.> I wish it had won over jq because JMESPath is a spec with multiple implementations and a test suite where jq is... well jq and languages have bindings not independent implementations.
jq has multiple implementations too! In Go, Rust, Java, and... in jq itself.
> jackson-jq aims to be a compatible jq implementation. However, not every feature is available; some are intentionally omitted because thay are not relevant as a Java library; some may be incomplete, have bugs or are yet to be implemented.
Where JMESPath has fully compliant 1st party implementations in Python, Go, Lua, JS, PHP, Ruby, and Rust and fully compliant 3rd party implementations in C++, Java, .NET, Elixer, and TS.
Having a spec and a test suite means that a all valid JMESPath programs will work and work the same anywhere you use it. I think jq could get there but it doesn't seem to be the project's priority.
commit|[?author == `Tom Hudson`]Not going to outright say that node.js scripts are the worst thing ever (they're not), but out-of-the-box Python is totally underrated (except on MacOS where `urllib` fails with some opaque errors untill you run some random script to deal with certs)
JSON.parse(<data>).foo?.[0]?.bar
Basically just return the `bar` field of the the first element of `foo`, or None/undefined if it doesn't exist. import json
data = json.loads('<data>')
bar = None
if foo:=data.get('foo'):
bar = foo[0].bar
print(bar)
If you can't be sure to get a dict, another type-check would be necessary. If you read from a file or file-like-object (like sys.stdin), json.load should be used.Even with that bias though, I have to admit that it's awful for typical command line script stuff.
Dealing with async and streams and stuff for parsing csv files is miserable (I just wrote some stuff to parse and process hundreds of gigs of files in node, and it wasn't fun).
Python is the right tool for that job IMHO.
Also, weirdly, maybe golang? I just came across this [1] and it has one of my eyebrows cocked.
Nushell[1] also seems like a promising alternative, but I haven’t had a chance to play with it yet.
Because its not.
Powershell is very nice as a glue language for .NET components, and its better as a general purpose shell/scripting language than the old DOS-inspired Windows Command Prompt, for sure.
I just can't stand title case, and Microsoft/.net absolutely love it. Everything in power shell is DoSomethingLikeThis.
Powershell is a great piece of tech that I just can't use because I'm old and grumpy and like snake or kebob casing.
I've never tried it on Linux though, so maybe it's different there?
Get-ChildItem == gciFor implementors case-insensitivity means the need for full Unicode support is urgent, while Unicode canonical equivalence does not often make the need for full Unicode support urgent. In practice one often sees case-insensitivity for ASCII, and later when full Unicode support is added you either have to have a backwards compatibility break or new functions/operators/whatever to support Unicode case insensitivity.
For users case-insensitivity can be surprising.
For code reviewers having to constantly be on the lookup for accidental symbol aliasing via case insensitivity is a real pain.
Just say no to case insensitivity.
https://stackoverflow.com/questions/33936074/decode-powershe...
Def. seen jq thrown into sed/awk scripts where a readable programming language was the right move. People spend hrs finding the right syntax to these things ~ not always well spent.
For my toolbox I include jq, gron, miller, VisiData, in addition to classics like sed, awk, and perl.
- https://github.com/saulpw/visidata - http://visidata.org/
Also there is a great introduction: - https://jsvine.github.io/intro-to-visidata/ "Intro to VisiData Tutorial" by Jeremy Singer-Vine
For simple things like navigating down one key, or one array entry, I know by heart, and it's incredibly useful. But anything more complicated, and I'm too lazy to lookup the documentation.
jq will fall into the bucket along with sed/awk of "tools I once wished to become an expert on, but will never do so because ChatGPT came along".
Would also put regex into that bucket, but they're so ubiquitous that I've already learned regexes. I wonder if the new wave of coders learning coding via ChatGPT will think of regexes the same way I think of sed/awk.
I've also found that learning by "ask ChatGPT, paste, verify" is so much faster and more fun than banging my head against concrete to deeply read documentation to reason about something new.
I've started doing this for new programming languages and frameworks as well, and it shortens the learning curve from months down to days.
It might be worthwhile to just learn how jq works. At the end of the day, you need to learn some language to parse json. I hate DSLs too, but I cannot think of anything as useful and concise as jq.
> but that ends up being at least as nasty as the JQ script
That's exaxtly why jq is so nice. Nice alternatives just don't exist
Putting that frustration on jq seems like a case of transference.
Of course not, but compared to every alternative today, jq is eons better than everything else. It's conciseness, ease of use, ease of learning all make it awesome. So as of right now, it is the nicest thing to use by far.
Personally though, I don't think I do wish for better. Jq is missing nothing that I want.
Write a simple Python script, parse JSON into native objects, manipulate those objects as desired with standard Python code, then serialize back into JSON if necessary. Voila, you have a readable, maintainable, straightforward solution, and the only dependency (the Python interpreter) is already preinstalled on almost every modern system.
Sure, you may need a few more lines of code than what would be possible with a tailor-made DSL like jq, but this isn't code golf. Good code targets humans, not "least possible number of bytes, arranged in the cleverest possible way".
jQ integrates very nicely into bash script. Especially in between pipes a short&simple jq-snippet can work wonders for readability of the overall script.
On the other hand, if the bash script becomes too complex it may be a good idea to replace the entire bash script with python (instead of just the json-parsing-part)
Many of them are not short and simple though. And each time you do a some transformation, you pretty much need to go in/out of jq at each step of it want to make some decisions or get multiple types of results without processing the original multiple times.
... if the reader happens to be familiar with the niche language "jq".
Otherwise, you may as well have put some Akkadian cuneiform in there.
jq seems slightly better than those...
Eh. Linux/Unix has always had an affinity for DSLs and mini-languages. If you're willing to work with bash, sed, awk, perl, lex, yacc, bc/dc etc. jq doesn't seem like it should cause too much consternation.
Second to this, I've mostly used jq to look at OpenAPI/swagger files, again just doing one-off tasks, such as listing all api routes, listing similarly named schemas, etc.
From what I've seen in the companies I've worked for, this is fairly consistent, but naturally I can't speak for everyone's use-cases. At the end of the day, I don't think most people use jq in places where readable or maintainable would be most appropriate.
https://gitlab.com/nbdkit/libnbd/-/blob/master/info/info-jso... https://gitlab.com/nbdkit/libnbd/-/blob/master/info/info-map... https://gitlab.com/nbdkit/nbdkit/-/blob/master/tests/test-ex... https://gitlab.com/nbdkit/nbdkit/-/blob/master/tests/test-ta...
(I picked a few completely at random from dozens and dozens)
It's beauty is the simplicity and portability.
Python is often not installed in server environments unless it's a runtime environment for Python.
Want to use a non standard library? Now your coworkers are suddenly in Python dependency hell. Better hope anyone else that wants to use this is either familiar with the ecosystem, or just happens to have an identical runtime environment as you.
Or someone could just curl/apt/dnf a jq binary to use your 3 line query, instead of maintaining all of this + 200 lines of Python.
(?:[A-Z][a-z]+_?(\d+))
Then I don't know what to tell you. Do you think that's too complex and should be a python script too? I don't think so. It looks complex, but if you just learn it, it's easier than a 'simple' script to do the same thing.I'd argue it's good code if you don't have to sift through lines of boilerplate to do something so trivial in jq or regex syntax.
echo '[{"name":"_skip"},{"name":"alpha"},{"name":"_other"}' | \
jq '[ .[] | select(.name|test("^_.*")|not) | . ]'
The same is roughly true for JMESPath, also, although at least it does actually try to allow projections and some limited functions $[?(@.name =~ /[^_].*/)]And, while whatever is powering https://jsonpath.com/ does honor your syntax, albeit with an absolutely useless result:
[
{
"name": -1
},
{
"name": -1
},
{
"name": -1
}
]
I found that `pip install jsonpath-ng` does not accept it nor mention it <https://github.com/h2non/jsonpath-ng/tree/v1.5.3?tab=readme-...> so I think it's out on the bleeding edge or somethingIt is also a single executable, written in clojure and fast. Among other niceties, you don't have to learn any DSL in this case -- at least not if you already know clojure!
Some people are weird and awe at the ellegance of piping 8 obscure commands, but if I'm given this shit and have to keep it working, I'm rewriting it on the spot.
Sometimes less general tools are nice. If they fit the problem space well, they can be very expressive without feeling unwieldy. And in some contexts reducing the power/expressivity is actually a good thing (e.g. not using a C interpreter to make your program and your config file use the same 'language')
Yeah if you don't like jq you likely won't like Regex, xpath, etc. Any syntax that is incredibly terse and complex.
Like Regex though, jq is too powerful to ignore and many times the best tool to use.
echo 12345678901234567890|jq
12345678901234567000For many people regexes are as bad as the jq queries… and vice versa. I would not recommend to write python script instead of regexp, but indeed it may work the same for small data and be more readable.
I love régex and been mastering it since 1999. So much that in 2013 I used it in production to parse binary protocol with dynamic sized fields. I believe the project is still talking 10k plus devices. Google must’ve just released protocol buffers… I would love to finally see regexes which can work over custom flow of objects and also on trees.
I also loved XPath which is very powerful and very comprehensible, then there is CSS1/2/3 which are again for queries to structures tree like data.
The prospect of now learning jq does not appeal me that much even though I appreciate its ingenuity. I may recommend it to dev/ops colleagues now and then, but for me this syntax is a lot of additional cognitive pressure which does not necessarily pay up. Of course if there is large amount of JSON data - it is the Swiss knife.
But nowadays I’ll likely use some LLm to generate the jq query for me. Also would joke with my bash-diehard colleagues who would love one more DSL…
That's almost certainly because both have pervasive generators/backtracking.
Whereas a few more modern shells have awk, sed and jq capabilities baked into the shell language itself. So you don’t need to mentally jump hoops every time you need to parse a different type of structured data.
It’s a bit like how you wouldn’t run an embedded Javascript or Perl engine inside your C#, Java or Go code base just to parse a JSON file. Instead you’d use your languages native JSON parsing tools and control structures to query that JSON file.
Likewise, the only reason jq exists is because Bash is useless and parsing anything beyond lists of bytes. If Bash supported JSON natively, like Powershell does (and to be clear, I’m not a fan of Powershell but for whole different reasons) then there would be literally no need for jq.
jq is way too much for what I need. I hacked together a filter in C to reformat JSON and I like it better than every JSON library/utility I have tried. For simple reformatting, jq is slow and brittle by comparison. Also, I can extract JSON from web pages and other mixed input. All the JSON utilities I have tried expect perfectly-formed JSON and nothing else.
Otherwise, its fine if you try to keep the thought "everything is a 'filter' or a composition of filters, and a 'filter' is a function that either maps, flatMaps or filters things" in your mind at all times
I love jq as a power tool and have the same challenges. I think the best path would have been for JavaScript to adopt something akin to JsonPath, although I more often reach to jq out of familiarity than use it in kubectl.
It does not work the same way as something like parsing an object and manipulating it in python. It is a query language. You are building up a result not manipulating objects.
Definitely unintuitive if you are coming from a programming language. Once learned it makes a lot more sense and is even preferable depending on your needs.
If I’m working on a Python script which has some jq embedded in it, then these problems probably exist:
- My editor will only syntax colour the Python, and treat jq code as a uniform string with no structure
- My linter will only consider Python problems, not jq problems
- My compiler, which is able to show parsing errors at compile time rather than runtime, will not give me any parsing errors for jq until execution hits it (yes, Python has a compilation step)
- jq error messages that show a line number will give me a relative line number for the jq code, rather than the real line number for where that code lives in the Python file
- My debugger will only let me pause and inspect Python, and treat the jq execution as a black box of I/O
I’m discussing this as a jq problem, but this happens far more commonly with SQL inside any host language. No wonder ORMs are so popular: their value isn’t just about hiding/abstracting SQL, it’s about wrangling SQL as a secondary language inside a different primary one.
- Microsoft’s LINQ for C#
- Webdev-focused IDEs which aim to correctly handle HTML and Javascript inside server-side languages (e.g. PHP)
- what else?
Any feedback is welcome !
Thank you so much for compiling it!
A related question for you and anyone else into this kind of tooling: if you had to automate some structural edits across a codebase that contains a wide range of popular languages (say: C++, C#, Java, Ruby, Python), and you had to do it with a single tool, which tool would you use?
You can get started with the Community Edition for free, it includes our JQ implementation (we call it kJQ - https://docs.kpow.io/features/data-inspect/kjq-filters)
Just shout if you need any help.
[1]: https://kislyuk.github.io/yq/
[2]: https://github.com/TomWright/dasel
[3]: https://hclq.sh/
https://github.com/tomnomnom/gron
I've been using `jq` for years and I'm always able to cobble together what I need, but I have yet to find it intuitive and I'm rarely able to arrive at a solution of any complexity without spending a lot of time reading its documentation. I wish I found it easier to use. :-(
But ChatGPT has genuinely solved my suffering writing jq, it does a pretty good job. It even almost replaces gron, if you feed it an exmaple json and ask for jq, it gives you something. It usually needs a little adjusting but it gets me 90% of the way there and saves me a bit of time.
I rarely use it for much else but its a jq winner :)
jq -r 'paths(scalars) as $p | getpath($p) | "\($p|join(".")) = \(.)"'
See elsewhere in this subthread for a full gron implementation in jq.[0] https://www.parkersoftware.com/blog/stop-using-simply-in-tec...
grep -A1 foo | grep -B1 bar
Will find a line with "foo" followed by a line with "bar" and emit both. Of course, it will also find a single line with both "foo" and "bar", so it's not perfect. This is a quick and dirty solution. Beyond that, break out sed and awk, or maybe the Practical Extraction and Report Language... it's really good at that stuff.The problem here is literally that someone hardcoded "IT'S ALWAYS LINEFEED" into an algorithm that could work equally well with any record separator character -- in fact probably with any record separator regex. I notice there's now `grep -z` which is one small step towards sanity... but the fully general problem is so easy and so useful to solve it's exasperating.
I guess I should stop complaining and submit a patch to grep to add a `--dont-use-linefeed-instead-use <arg>` option already.
E.g.,
# gron
gron "https://api.github.com/repos/tomnomnom/gron/commits?per_page=1" |
fgrep commit.author
json[0].commit.author = {};
json[0].commit.author.date = "2016-07-02T10:51:21Z";
json[0].commit.author.email = "mail@tomnomnom.com";
json[0].commit.author.name = "Tom Hudson";
# jq
curl -L "https://api.github.com/repos/tomnomnom/gron/commits?per_page=1" |
jq -r 'paths(scalars) as $p | getpath($p) | "\($p|join(".")|select(contains("commit.author"))) = \(.)"'
0.commit.author.name = Tom Hudson
0.commit.author.email = mail@tomnomnom.com
0.commit.author.date = 2022-04-13T14:23:37Z
# jq with grep outside jq
curl -L "https://api.github.com/repos/tomnomnom/gron/commits?per_page=1" |
jq -r 'paths(scalars) as $p | getpath($p) | "\($p|join(".")) = \(.)"' |
fgrep commit.author
0.commit.author.name = Tom Hudson
0.commit.author.email = mail@tomnomnom.com
0.commit.author.date = 2022-04-13T14:23:37Z
With just a bit more work you can get it to output valid gron, and even to parse valid gron.the biggest usecase for me is taking some csv, toml, xml, whatever and converting that to json so I can pipe to jq
It's inspired by XPath so it's very familiar instead of a complete new DSL. The killer feature imo is the recursive key lookup so you can write `people..address` and it'll find all "address" keys that descend from "people" anywhere in the JSON. It's by far my favorite parsing language for JSON and I wrote an introduction blog on how to use it in JSON dataset parsing [2] :)
You can even `gron | grep | sed | gron -u`.
Awesome tool thanks for sharing.
It supports everything gron supports, but 50x faster.
I'll also share that I've been using `curl cheat.sh/jq` (cheat.sh in general is a great resource) for years.
Although now I'd probably use something like chatgpt.
I have some code around for yq instead that keep breaking because yq keeps improving in non backward compatible ways (I didn't investigate how often yq introduced backwards incompatible changes, but the issue affected me several times in unrelated places, CI scripts or whatnot, that by their nature end up running with different versions of base tooling and update them at various pace)
I was always thus grateful to the great wisdom of the jq maintainers for their understanding of the importance of backwards compatibility.
I hope this announcement doesn't mean that this stability was just an accidental side product of stagnation and that once stagnation is "fixed" it will be done at the expense of stability.
- https://kislyuk.github.io/yq/
- https://github.com/kyle-long/yq#yq
- https://github.com/up1/yq-1#yq-command-line-yamlxml-processo...
- https://github.com/onixspot/yq-2#yq-command-line-yamlxmltoml...
[1]: I personally think that JSON is plenty readable, but a lot of people seem to disagree.
Its error reporting is also clang-vs-gcc level wizardry, and I often use it to get a helpful message instead of "ENOWORKY" from jq (I haven't tried 1.7 yet, so it could be better for all I know)
I'd personally use that, but in the context of sharing scripts and snippets with colleagues, the strength of the incumbent `jq` is that we can all assume everybody will have it installed on their machine.
Is there a way to get jq version inside the script?
This is the first I'm hearing of gron, but adding here for completeness sake. Meanwhile, JSON seems to be becoming a standard for CLI tools. Ideal scenario would be if every CLI tool has a --json flag or something similar, so that jc is not needed anymore.
Huge fan, I use it all the time.
It's really awesome how the community pulled together and helped us recruit new maintainers to revive the project. Special thanks to, well, all involved, but especially @stedolan, @itchyny, and @owenthereal (all GitHub usernames).
> jq -n '{"a": 1, "b": {"c": 2, "d": 3}, "e": 4} | pick(.a, .b.c, .x)'
This is a godsend! Thanks to the contributors! <3
$ jq -n '{"a": 1, "b": {"c": 2, "d": 3}, "e": 4} | {a, e}'
{
"a": 1,
"e": 4
}Personally, when I test REST APIs, I use „restclient.el“ all the time which also comes with a great JQ integration („jq-set-var“ for example for deriving request variables from responses). For traversing larger responses I use „counsel-jq“ in a customized JSON mode: https://github.com/200ok-ch/counsel-jq
But I’ll give the major mode a try, too.
But at one point I started write long jq modules and while it was pretty straightforward, there are less people familiar with jq.
So I declared jq bankruptcy and rewrote it as a nodejs script. The rest of the team was relived
If you're writing long jq modules, you probably do want a different (faster, better) language.
Strangely, I also have ECMA-404 and RFC8259 open in other tabs. mostly annoyance with the occasional flashes of anger over number formats and duplicate keys.
First release of jq in 5 years - https://news.ycombinator.com/item?id=36951830 - Aug 2023 (27 comments)
jq -r \
'.apps.http.servers.srv1.routes[0]
| .match[0].header.[env.AUTH_USER_HEADER][0] = "$username"'
That's an error because you can't select an env var key '[env.AUTH_USER_HEADER]' in the middle of a chain like that, only immediately following a pipe: | .match[0].header | .[env.AUTH_USER_HEADER][0] = "$username"'
But then I need to preserve the parent object after the assignment and that pipe throws it out. Thankfully, parentheses fix that: | (.match[0].header | .[env.AUTH_USER_HEADER][0]) = "$username"'
After working out the gotchas, it's quite powerful, like regex, but a little clunky, though not nearly as much as regex.But isn't that what jq is and always has been? I mean,what led you to believe that an entirely unrelated pattern matching language would work?
> That's an error because you can't select an env var key '[env.AUTH_USER_HEADER]' in the middle of a chain like that, only immediately following a pipe:
> | .match[0].header | .[env.AUTH_USER_HEADER][0] = "$username"'
You can use `env.AUTH_USER_HEADER` as a key the way you wanted. The issue is that you had to write `... | .match[0].header[env.AUTH_USER_HEADER] ...` -- no `.` between "header" and the index operator!
This complaint is a fairly frequent one, so in fact we did "fix" this in 1.7! You can now write `.a.[0]` and it works.
I often use ripgrep to setup quick bash pipelines for rapid data analysis, would love to be able to use jq for that purpose. These days I am setting up scripts with simdjson but the cost of writing a script vs quickly setting up jq or ripgrep in a bash pipeline are orders of magnitude different.
It's a big part of our check evaluation infra at OpsLevel.
It's a line-editor, and allows me to query json documents ultra-fast. It optimizes for a really fast feedback.
Postman handles this especially poorly...
* Run directly from the command line
* Small interpreter ~1MB
* Compact language (for better or worse)
* Stable