Show HN: jql – Easier jq Alternative with a Lispy Syntax Written in Go
github.com
github.com
It inspired me to create an alternative to jq with more of a Lispy syntax, as I think the original is awesome but also fairly cryptic for anything more advanced than single field selection.
Overall it was a fun few-day project, and was also very educational in terms of writing a parser in Go. (I used goyacc before in the sql parser of OctoSQL[1], however, that one is copied from vitess, so I’ve never built one from scratch, only customised an existing one. It’s really pleasant overall and the code is very simple, so I encourage you to take a look[2].
I'd love to hear any feedback, comments or potential improvements you can think of.
Sorry for the name duplication! I thought I'm sharing the name only with the jira API jql.
EDIT: I'd love to see some benchmarks. I expect your jql to greatly outperform mine and jq, considering it's written in Rust!
Is this understandable? "You can see that elem is the most used function, and in fact it's what you'll usually be using when munging data, so there's a shortcut. If you put a value in function name position, it implicitly converts it to an elem."
EDIT: Fixed the order.
I scratched a similar itch by creating Jowl[1], which uses JavaScript one-liners with the Lodash library to transform JSON.
Like you, I wanted syntax that was more familiar to me than jq's, and like you, I wanted to use a language that was built for data structure transformation.
However, I offloaded all of the actual hard work to an existing JavaScript runtime, and to an existing library for data structure transformation, meaning I could punt on the parser and language design; the hard parts that you've taken on. I have this bookmarked to read through in more detail later this week.
My approach was caused by the fact that I wanted to have a minimal language which only contains what really is necessary and is as regular as possible, to be very simple as a result.
As an old Lisper, I did get a wry chuckle out of the fact that you managed to avoid Lisp for a tool that uses Lispy syntax. :-)
Yup, I felt the irony myself creating it, as I actually love clojure!
But I'm most comfortable with Go, and it results with a single binary for any of the major platforms with very quick startup time (as illustrated by the small benchmark), so those were huge advantages.
(elem "countries" (elem (keys) (elem "name")))
I think part of the problem is, IIUC, this is constructing a pipeline, like (->) in Clojure, but is using function call syntax rather than pipeline syntax, and also mixing in a function call with (keys). I feel like it's a very non-Lispy language masquerading as Lisp by wearing parens.Instead, if you wrote it as real Lisp function calls, and provided a real pipeline macro, it would be much more Lispy and, I think, much easier to understand. For example, instead of:
(elem "countries" (elem (keys) (elem "name")))
Either of these alternative syntaxes: (-> "countries" (key "name"))
(key "name" ("countries"))
Would return the same result: ["Poland",
"United States",
"Germany"]
While being more concise and more Lispy.I also wouldn't mind using an anaphoric 'mapcar, something like:
(-> "countries" (map-> (key "name"))
Otherwise it may be unclear which functions implicitly map across entries and which don't.My two cents. Thanks for sharing your work.
Have you checked out the type cheatsheet? (keys) is a function call!
The function you pass as last to elem is kind of a continuation. It's a function which transforms the output before returning.
All you're doing writing a jql query is composing one big function.
With values in function call position, it's just that (keys) gets evaluated to a value (the keys of the map) and that is used to index the json. That's why keys wouldn't work, you need the parentheses to make it a function call.
This may sound complex at first but I find it gets intuitive quickly.
EDIT: Check out this comment thread on reddit as I think the commenter may have had similar problems: https://www.reddit.com/r/golang/comments/ehnsz5/jql_json_que...
I'm afraid that that does not help me understand how (elem) works at all. From my perspective, I don't need a "type cheatsheet," I need elem's docstring that explains what arguments it expects and what it returns. I'm not even thinking about value types yet, and I don't know what grammar that cheatsheet is written in, nor why I would need to read one to be able to use what seems like the most basic function jql has, the elem function.
To write clear, useful documentation, you need to put yourself in the shoes of a user who has never seen your project before, who has not walked the miles you have to arrive at the solutions you have, and take them quickly to the same destination.
The readme, in general, is another issue. It's...not ideal.
The first part tries to sound cute with stuff like, "Hey there!" and, "remember? That's explicitly not why we're here. But it aids understanding of the more complex examples, so stay with me just a little bit longer!".
But I'm not a little kid who's bored in math class, so writing like that turns me off quickly. Having to read through long prose like, "Ok, let's check it out now, but first things first, you have to install it:" instead of simply a heading titled "Install:" feels like I'm wasting my time, and I quickly skip ahead or move on to the next HN article.
What I'm really looking for is a short intro and a table of contents with sections like: Intro, Examples, Tutorial, Reference, FAQ.
You might think of it like this: the project is your baby, but just like in real life, no one thinks your baby's as cute as you do, and, no, we don't want to see the baby pictures. ;)
Everybody has different taste. I like README’s like that as they add an entertainment value. Overall I’ve got two kinds of feedback about the README: 1. I really like the project and really liked the readme, great fun to read! 2. The readme was very confusing on top of the already confusing query language.
And based on that I'm planning to keep it unchanged for now. I added the type cheatsheet to make case 2 at least a little bit better.
As for the quick examples, you can just scroll through the code blocks (as each one contains input with its corresponding output).
I'm thinking of adding another one which shows how standard map/filter notation maps to jql queries.
Anyways, thanks for the feedback again, I do appreciate it!
In practice, thanks to this, the resulting queries end up matching the input JSON structure very well.
(elem x f)
is sugar for (filter-on f (elem x))
or so? (filter-on imaginary thing but hopefully obvious enough as pseudocode) $ jsonnet -e "(import 'test.json').countries[0]"
$ jsonnet -e "(import 'test.json').countries[0:2]"
$ jsonnet -e "[x.name for x in (import 'test.json').countries]"
$ jsonnet -e "std.objectFields((import 'test.json').countries[0])"
With that in mind, I never bothered to learn how to use a tool like jq, rq, jql, etc. open test.json -> [ countries ]
open test.json -> [[ /countries/0 ]]
open test.json -> [ countries ] -> [ 0 2 ]
Single brackets only return the next level deep but can retrieve multiple items (and negative integers to count from the end of an array). Double brackets are to specify a path. So [[/foo/bar]] is literally the same as [foo]->[bar]Lastly the 4th example:
open test.json -> [ countries ] -> foreach c { echo $c[name] }
Unfortunately this one doesn't output it in JSON. If you needed that you could chain another command to reformat it: open test.json -> [ countries ] -> foreach c { echo $c[name] } -> cast str -> format json
This works because murex can auto-convert between lists, JSON, YAML, CSV and a few other structured formats. However it's fair to say it is a lot more verbose than jsonnet and jql in that last example.Github repo: https://github.com/lmorg/murex
I've been using this as my primary shell for a few years now and, like yourself, I've never bothered to learn jq because of that.
NB all of the above example are running from inside the murex interactive command line (like bash). If you wanted to run it like jq/jql then you'd need to do the same sort of thing as bash:
murex -c 'open test.json -> [ countries ]'I haven't used it yet, but the syntax seems simpler than some.
I never before wanted to write object oriented JSON, but now my mind has been opened.
For those who haven't yet seen it/opened the link, it's 20% project that describes itself as:
> A data templating language for app and tool developers
- Generate config data
- Side-effect free
- Organize, simplify, unify
- Manage sprawling config
> A simple extension of JSON
- Open source (Apache 2.0)
- Familiar syntax
- Reformatter, linter, editor & IDE integrations
- Formally specified
For an Open Source project of mine (https://github.com/buildbarn, a remote build cluster implementation for the Bazel build system), we are using go-jsonnet as the config file parser library. This means that people can use JSON, but if they need something that’s more flexible, they can add Jsonnet statements immediately. It’s also possible to have that as a preprocessing stage (init scripts), but in a containerized world it’s easier to have it embedded.
One of the nice things about doing that is that I also don’t need to write code to load TLS certificate data from disk. People can either embed their secrets directly into the config file, or they can use Jsonnet’s ‘importstr’ keyword to load separate PEM files from disk.
Somewhat related to that: Jsonnet also allows people to write libraries for constructing more complex JSON files. For example, the Grafana folks have released a library (https://github.com/grafana/grafonnet-lib) that allows you to easily build Grafana dashboards from code.
That sounds awesome. I'd expect it's worth it just to have self and $.
And not only that, jsonnet can also produce yaml and ini files. I might try using it to configure my whole machine...
> TODO: This page still needs to be written.
Looks like it is not ready for prime time yet... (Not to mention that it seems to have no support for stdin input, multiple input files, and json lines format - which are my main uses cases for jq)
Though from the tutorial I don't really understand how I can use it to transform data, other than converting between formats.
I'm using https://github.com/borkdude/babashka with with https://github.com/borkdude/jet and a few aliases along with curl for the same effect.
The syntax of jq is so hard to remember a familiar lisp always beats it :)
JSONata remains pretty obscure but it's been my favorite JSON query/mangling DSL for a while now: https://github.com/jsonata-js/jsonata.
jmespath is also quite natural, but I think it was hurt by reference implementation in Python, meaning it's quite slow. http://jmespath.org/
I think it can get a little weird as far as error handling goes, but I feel it's been pretty approachable.
It even features a complete rewrite of the actual jq in Go.
I had no idea about jq-mode for Emacs, it even has org-babel support. So cool!
echo '' | fzf --print-query --preview-window wrap --preview 'cat test.json | jql {q}'
Something close to: echo '' | fzf --print-query --preview-window wrap --preview 'cat test.json | jq {q}' echo '' | fzf --print-query --preview-window wrap --preview 'cat test.json | jql {q}'
Trivial note. echo instead of echo '' or echo "". $> [[ "$(echo)" == "$(echo '')" ]] && echo $? || echo $?
$> 0 $> [[ "$(echo)" == "$(echo '')" ]] && echo $? || echo $?
$> 0
Trivial note. `; echo $?` instead of `&& echo $? || echo $?` $> [[ "$(echo)" == "$(echo '')" ]]; echo $?
$> 0Highlights - interactive mode - Use full power of JavaScript.
$ curl ... | fx '.filter(x => x.startsWith("a"))'
- Access all lodash (or ramda, etc) methods by using .fxrc file.
$ curl ... | fx '_.groupBy("commit.committer.name")' '_.mapValues(_.size)'
echo '' | fzf --print-query --preview-window wrap --preview 'cat test.json | jql {q}'So far I'm planning to just fire up Hy the next time I need anything in this vein. Since it's on Python, both JSON parsing/encoding and functional stuff should be there out of the box. Alternatively, I still can try using ClojureScript with Lumo. But I don't expect too fast startup with either of them, compared to, say, Fennel/Lua.
BTW, there's a Lisp that translates to Go: https://github.com/jcla1/gisp. Though not too popular, and I've heard an opinion that it's toy-like in its capabilities.
Thanks for the inspiration though!
Though you can surely put a oneliner together which achieves the same effect using fzf, as per: https://paweldu.dev/posts/fzf-live-repl/
echo '' | fzf --print-query --preview-window wrap --preview 'cat test.json | jql {q}'https://youtu.be/PS_9pyIASvQ [Serious Programming with jq?! A practical and Purely Functional Programming Language!]
Would be an error.
Expression evaluation ought to be inside out.
It is, it's just that everything returns a function! The topmost function will in the end be called with the json as an argument.
Check out these parts of the README: https://github.com/cube2222/jql#attention https://github.com/cube2222/jql#type-cheatsheet
(elem 0) returns a function which takes a JSON array as input and returns its first element.
It's just that the query ends up having a type JSON -> JSON, not just JSON.
(elem "countries" (elem 0))
Inside out elem 0 has to return a function since it doesn't have enough to do anything. I'm not sure why 0 couldn't be a key necessitating a different function to differentiate but whatever.Elem countries has to return json and one can imagine using the json in the first position like a map in clojure and taking a fn in following position. This would mean that on net it would be exactly the same as.
(elem 0 (elem countries))Yours is more akin to (elem 0 (elem "countries" json))
and you can see that the hierarchy of the query is now inverted in respect to the json. This is what I like about the continuation based approach, a big query will be readable because it fits the data very well.
(->> JSON
(elem "countries")
first)To me it's basically the same readability wise with proper indentation.
Though with deeply nested data and a lot of map functions I think you'd have a lot of unnecessary (->> JSON ...). But those can probably be eliminated with another macro.
Also, there's the pipe function in jql which basically lets you do this.
Programming with combinators over functions is a thing also.
(elem (elem input "countries") 0)From a quick look on the repo, this seems so much simpler. Thanks!!!
It just evaluates jql continuously while you can write and edit your query.
jq and lisp?
In all seriousness, thank you for writing this. I've had trouble wrapping my small brain around some of jq's more advanced syntax, so being able to look at things from a different direction seems like a great idea to me.
Let me know if it was understandable/confusing overall!
There's an issue for at least adding trailing and starting parentheses automatically, so I'll try to add at least the trailing ones.
It also converts between yaml and json with ease for writing and reading
Also, jq supports modules. which is awesome!
[0]https://marketplace.visualstudio.com/items?itemName=ldd-vs-c...
Perhaps I would want some implicit pipe or threading at the top level (which could be turned off with some option for scripts I guess).
Another helpful feature could be adding the magic closing paren: ] which will close all open parens to the top level. E.g. these two lines are equivalent:
(foo (bar (baz) whizz))
(foo (bar (baz) whizz]
A completely different (and less powerful) paradigm I like is one of match-format where the pattern looks a lot like what it’s trying to match. In this case most functions and fancy transformations are impossible. So you could write something like this to extract the list of countries: j '{countries: %%}'
Or to extract each country’s name: j '{countries: [ .*, { name: %% }, .*]}'
(Presumably you would have syntax for trying to match against each array element). (I also would make commas optional but I’m trying to make this syntax obvious).Or to list the keys of the 0th element:
j '{countries: [ { %%: . }, .* ]}'
Or to get name and population: j '{countries: [ .*, { name: %name, population: %pop }]}'
The output might look like: { "name": "Poland", "population": "38000000" }
————————An alternative query tool which I haven’t really seen before would be one which is interactive rather than repeatedly run: you pipe your json (or whatever) into it and it spins up some kind of (curses) gui shoeing you your data with a nice interface to interactively write your query and see its results. It could finish by outputting a command line to run the query non-interactively.
The closest thing to this I know of is a small programming by example demo that came out of Microsoft research. The goal was to figure out how to extract some desired data from some big blob of json (or maybe html, I don’t remember). The user would see the json in some web gui and could click on the data they wanted and the app would try to figure out the “most obvious” query to get it. If they clicked on the country name, I think it would probably suggest a query for all the countries’ names (but maybe that would be the second option). I think it would also try to be clever about grouping, so if the user picked name and population it would hopefully figure out that they should come in pairs rather than being two independent lists that might not be the same length.
There's already an issue to automatically add trailing parens and I definitely think it's a good idea, as balancing parens really is annoying in-shell. I didn't think of it before.
I decided that adding interactivity would be too much work, but you could definitely put together a oneliner (or alias) which achieves that using fzf: https://paweldu.dev/posts/fzf-live-repl/
Overall I encourage you to try write a tool that has the syntax you described, as it sounds plausible! It kinda looks a little bit like graphql, so maybe that would fit?
I'll think about it.
echo '' | fzf --print-query --preview-window wrap --preview 'cat test.json | jql {q}'