Q: A faster re-implementaiton of jq written in Reason Native/OCaml
github.com
github.com
https://github.com/fiatjaf/awesome-jq
https://github.com/TomConlin/json2xpath
https://github.com/antonmedv/fx
https://github.com/fiatjaf/jiq
https://github.com/jmespath/jp
https://github.com/cube2222/jql
https://github.com/borkdude/jet
https://github.com/jzelinskie/faq
https://github.com/dflemstr/rq
Personally I think that next time I might just fire up Hy and use its functional capabilities.
https://docs.microsoft.com/en-us/powershell/module/microsoft...
It is not nearly as expressive as jq, but it is faster for my use cases (written in golang).
Personally I'd prefer Fennel, which is on Lua and thus a whole lot faster, especially in regard to the startup time—but as I noted in a thread on Fennel, Lua's omission of a proper ‘null’ makes it awkward to handle exchange and transformations of data from third parties. And, since I'm likely to fiddle with the queries for some time, startup delay is less important here.
https://github.com/borkdude/babashka
https://news.ycombinator.com/item?id=24353476
Aside: another nice tool I recently discovered for working with JSON and YML, doing conversion and diffs (especially helpful for generated files):
> Make JSON greppable!
> gron[1] transforms JSON into discrete assignments to make it easier to grep for what you want and see the absolute 'path' to it. It eases the exploration of APIs that return large blobs of JSON but have terrible documentation.
▶ gron "https://api.github.com/repos/tomnomnom/gron/commits?per_page=1" | fgrep "commit.author"
json[0].commit.author = {};
json[0].commit.author.date = "2016-07-02T10:51:21Z";
json[0].commit.author.email = "mail@tomnomnom.com";
json[0].commit.author.name = "Tom Hudson";
[1] https://github.com/tomnomnom/gron2) doesn't accept values that jq accepts
% time jq -r '[expression]' < parcels | wc
365 1454 7978
jq -r < parcels 1.39s user 0.00s system 99% cpu 1.390 total
wc 0.00s user 0.00s system 0% cpu 1.390 total
% time ~/.yarn/bin/q '[expression]' parcels | wc
q: internal error, uncaught exception:
Yojson.Json_error("Line 56, bytes -1-32:\nJunk after end
of JSON value: '{\n \"OBJECTID\": 155303,\n \"BOOK\"'")When using jq, I can do a lot of things:
aws s3 cp s3://bucket/file.json.gz - | zcat | head | jq .field | sortThat's how Go's static file web server works. It serves streams, but if you happen to io.Copy to that stream with something that is also an ∗os.File on Linux, it can use the sendfile call in the kernel instead. (A downside of making it so transparent is that if you wrap that stream with something you may not realize that you've wrecked the optimization because it no longer unwraps to an ∗os.File but whatever your wrapper is, but, well, nothing's perfect.)
jq .field <(aws s3 cp s3://bucket/file.json.gz - | zcat | head) | sort
which is more annoying to type but works.So, this should work :-)
% cat /proc/self/cmdline <(echo $SHELL) | tr '\0' ' '
cat /proc/self/cmdline /proc/self/fd/11 /bin/zshIt won't work in dash though, and you should not use this in a shell that targets POSIX.
It's not hard to fix things like this, but it exemplifies a lack of familiarity with the Unix command line. There are an enormous number of tools out there that only exist because people don't know how to chain together basic 1970s Unix text-processing tools in a pipeline.
That's not to say Go's decisions to toss some established practices are "wise" or "sagely", just that broad acceptance is not a criteria they seemed concerned with. Which is fine.
>they think they know better.
It's safe to say Rob Pike is not clueless or without experience in unix tooling. You should listen to some of his experiences and thoughts with designing Go [0]. I don't always agree with him, [but it's very baseless to suggest he makes decisions on the grounds that they were his, not they have merrit.]
Edit to clarify: [He makes decisions on merrit over authority]
I'm (perhaps unfairly) uninterested in writing out all the details, but “they think they know better” is because I see Go as someone's attempt to update C to the modern world without considering the lessons of any of the languages developed in the meantime. And because of the weird dogmatic wars about generics, modules, and error handling.
I'm actually surprised by this; I would have expected Pike to go with single-letter options only.
> Rob 'Commander' Pike
> Apr 2, 2013, 6:50:36 AM
> to rog, John Jeffery, golan...@googlegroups.com
> As the author of the flag package, I can explain. It's loosely based on Google's flag package, although greatly simplified (and I mean greatly). I wanted a single, straightforward syntax for flags, nothing more, nothing less.
> -rob
Myths like "everything is a file" or file descriptor is complete bollocks, mostly retconned recently with Linuxisms. Other than pipes, IPC on Unix systems did not involve files or file descriptors. The socket api dates to the early 80s and even it couldn't follow along with its weird ioctls. Why are things put in /usr/local anyway? Why is /usr even a thing? There's a history there, but these days I don't seem much of anything go into /usr/local on most Linux distributions.
It's also ironic to drag OS X into a discussion of Unix, because if there was one system to break with Unix tradition (for the best in some ways) -- no X11, launchd, a multifork FS, weird semantics to implement time machine, a completely non-POSIX low-level API, etc, that would be it.
All this shit has been reinvented multiple times, the user-mode API on Linux has had more churn than Windows -- which never subscribed to a tradition. There's no issue of lack of familiarity here, the original Unix system meant to run on a PDP-11 minicomputer only meets modern needs in an idealized fantasy-land. Meanwhile, worse is better has been chugging along for 50 years while people try to meet their needs.
XQuartz if you want it
> completely non-POSIX low-level API
macOS has a POSIX layer.
There are X server implementations for Windows, Android, AmigaOS, Windows CE!!, etc... I don't think this is relevant.
> macOS has a POSIX layer. So do many systems, again including Windows in varying forms through the years. I think the salient issue is that BSD UNIX and "tradition" are conflicting. The point of the original CMU Mach project was to replace the BSD monolith kernel.
My understanding is that Windows has always had a very strong tradition of backwards compatibility. Even to the point of making prior bugs that vendors rely on still function the same way for them (i.e. detect if it's e.g. Photoshop requesting buggy API, serve them the buggy code path and everyone else the fixed one).
That's just as much a tradition as "we should implement this with file semantics because that's traditionally how our OS has exposed functionality".
You have always been able to customize Homebrew to install at a custom prefix, e.g. ~/brew. It’s just that, when you do that, and then install one of the casks or bottles for “heavy” POSIX software like Calibre or TeX, that cask/bottle is going to pollute /usr/local with files anyway, but those files will be symlinks from /usr/local to the Homebrew cellar sitting in your home directory, which is ridiculous both in the multiuser usability sense, and in the traditional UNIX “what if a boot script you installed, relies on its daemon being available in /usr/local, which is symlinked to /home, but /home isn’t mounted yet, because it’s an NFS automount?” sense. (Which still applies/works in macOS, even if the Server.app interface for setting it up is gone!)
The real ridiculous thing, IMHO, is that Homebrew doesn’t install stuff into /usr, like a regular package manager. But due to macOS considering /usr part of its secure/immutable OS base-image, /usr is immutable when not in recovery mode.
I guess Homebrew could come up with its own cute little appellation — /usr/pkg or somesuch — but then you run into that other lovely little POSIXism where every application has its own way of calculating a PATH, such that you’d need to add that /usr/pkg directory to an unbounded number of little scripts here and there to make things truly work.
...and then try to build something entirely sensible like Postgres, but hours of fiddling with different XCode versions and compiler flags still lead to a dead end of errors, you're stuck because you're running an unsupported configuration.
I still don't understand how the PG bottles for Mojave can be built.
Homebrew just acknowledges that these external third-party binary distributions (casks) are going to make a mess of your /usr/local — because that's the prefix they've all settled on burning in at compile-time — and so Homebrew tries to at least make that mess into a managed mess.
And, if some other system is already managing /usr/local, but isn't expecting the results of these programs unpacking into there, it's going to be very upset and confused — again, regardless of whether or not you use Homebrew. So it'd be better for those other systems to just... not do that.
/usr/local isn't supposed to be managed. It's supposed to be the install prefix that's controlled by the local machine admin, rather than by the domain admin. Homebrew just happens to be a tool for automating local-admin installs of stuff.
/opt/homebrew would be a somewhat traditional place to put it.
> but then you run into that other lovely little POSIXism where every application has its own way of calculating a PATH, such that you’d need to add that /usr/pkg directory to an unbounded number of little scripts here and there to make things truly work.
What? You should be able to add it to the system PATH that's set for sessions and call it a day on a POSIX system. PATH is an environment variable and inherited. If MacOS is in the habit of overriding PATH on system scripts I have to imagine that's because they completely screwed it up at some point in the past. Generally, you just add it to your use session variables in whatever way your system supports (.profile, etc) if you want it for your user, or at a system level if you want it system wide (I could see maybe Apple making this hard).
The only times in over 20 years I've ever had to deal with PATH problems are when I ran stuff through cron, because it specifically clears the PATH. More recent systems just specify a default PATH in /etc/crontab for the traditional / and /usr bin and sbin dirs.
Maybe you're thinking of the shared library path loading? That should also be easily fixed.
Mac OS X had some of the sexiest ways to install and uninstall application software that we'd ever seen in any other platform at that time.
But that Apple stubbornly refused to include a useful package management system, was one of the most horrible oversights in computing history.
Not necessarily. Plenty of software uses relative paths that work regardless of prefix. Off the top of my head, Node.js is distributed in this way.
> you’d need to add that /usr/pkg directory to an unbounded number of little scripts here and there to make things truly work.
How so? Are there that many scripts that entirely replace the PATH environment variable? In Linux, I just include my system wide path additions in /etc/profile which will be set for every login. For things like cron jobs or service scripts, which don't inherit the environment of a login shell, you will need to source the profile or use absolute paths, but that's about the only caveat I can think of.
Fully agree with you, but oh well, most if not everything is available on Macports anyway.
> There are an enormous number of tools out there that only exist because people don't know how to chain together basic 1970s Unix text-processing tools in a pipeline.
Speed. A specialized tool you need often beats manually wrangling the dozen or so Unix tools you need to replace it, plus many Good Options are only available on the GNU/Linux coreutils and don't work on Macs (sed -i, my most common annoyance) or busybox.
Arguably that is why the original implementation of Perl was written. If I remember the story correctly, we can never know for sure whether, e.g., AWK would have sufficed, because the particular the job the author wrote Perl for as a contractor was confidential.
Are people using jq most concerned about speed, or are they more concerned about syntax.
JSON suffers a problem from which line-oriented untilities generally have immunity: a large enough and deeply nested JSON structure will choke or crash a program that tries to read all the data into memory at once, or even in large chunks. The process is resource-constrained as the size of the data increases. There are no limits placed on the size or depth of JSON files.
I use sed and tr for most simple JSON files. It is possible to overlfow the sed buffer but it rarely ever happens. sed is found everywhere and it's resource-friendly. Others might choose a program for speed or syntax but the issue of reliability is even more important to me. jq alone is not a reliable solution for any and all JSON. It can be overkill for simple json and resource-constrained for large, complex JSON.
https://stackoverflow.com/questions/59806699/json-to-csv-usi...
netstrings (https://cr.yp.to/proto/netstrings.txt) do not suffer from the same problem as JSON.
Yes, q is supposedly faster than jq. But it is exceedingly rare for me to ever have any performance problems with jq, especially since it’s essentially a one off utility I use occasionally, not as part of the hot loop of any workflow where performance matters.
The incompatibility is apparently due to the fact that jq is happy with a concatenation of JSON objects and q is not. For example {'foo':1}{'foo':2} as opposed to [{'foo':1},{'foo':2}]
I certantly didn't use it, but I see where it's useful, will implement it soon.
The comma operations means that the "filters" are duplicated, so instead of one json state you would pass two if there's one coma; and ofcourse any number of commas are allowed.
I definitely agree that reading from stdin is critical if I'll be able to use it. Don't take the criticism too hard though (especially the "author doesn't appreciate unix" stuff. Sometimes we can be such assholes to each other).
Nice work!
Judging is free and I didn't consider stdin as something to spend time on yet. Will do, now that some people raise it.
Thanks :D
echo '{"foo": "bar"}' | query-json ".foo" /dev/stdin # Cuts off partway through:
zcat atop.log.gz | atop -r /dev/stdin | less
# Works fine:
zcat atop.log.gz > atop && atop -r atop | less
That said, I agree with your experience that /dev/stdin usually works for programs that read a file straight in.> Normally, file only attempts to read and determine the type of argument files which stat(2) reports are ordinary files. This prevents problems, because reading special files may have peculiar consequences.
One example that comes to mind is /dev/urandom, which sucks randomy values out of the entropy pool (at least in Linux)—and the pool can be exhausted, or at least it could back in the day, not sure about now. Other possible cases are things in /proc (though unlikely), and particularly stuff like serial ports—where presumably reading could gobble data intended for some drivers or client software.
Two examples I can remember off the top of my head:
- Nix build scripts
- OpenMoko
query-json ".foo" <(echo '{"foo": "bar"}')Happy to rename it to qj instead, but the option of renaming the binary it's a good workaround.
He replied and thought about qj
Query CSV with SQL
I will try to bring it to the brenchmark, thanks for sharing
echo '' | fzf --print-query --preview-window wrap --preview 'cat test.json | jql {q}'
It is more verbose, like you get the size of something with array:size or map:size functions, so it is more readable
I am implementing it in Xidel 0.9.9+: http://www.videlibri.de/xidel.html
xmlstarlet supports XPath 1, but the W3C did not stop there. They made XPath 2 featuring variables, lists and regular expressions, XPath 3 featuring higher order function, and finally XPath 3.1 featuring JSON support.
For example,
echo '[{"a": 1}, {"a": 2, "b": 3}, {"c": 4}]' | xidel - -e '?*?a!(. * 100)'
will print 100 and 200.Or the same with a verbose syntax:
echo '[{"a": 1}, {"a": 2, "b": 3}, {"c": 4}]' | xidel - -e 'for $obj in array:flatten(.) return map:get($obj, "a") * 100' ~/xidel % time ./xidel - -e '?SitusAddress!(.)' < ~/parcels | wc
**** Processing: stdin:/// ****
29066 149317 903704
./xidel - -e '?SitusAddress!(.)' < ~/parcels 2.55s user 0.18s system 99% cpu 2.733 total
wc 0.01s user 0.00s system 0% cpu 2.733 total
~/xidel % time jq -r '.SitusAddress' < ~/parcels | wc
29066 149317 903704
jq -r '.SitusAddress' < ~/parcels 0.95s user 0.00s system 99% cpu 0.958 total
wc 0.00s user 0.01s system 1% cpu 0.957 total> The report shows that q is between 2x and 5x faster than jq in all operations tested and same speed (~1.1x) with huge files (> 100M).
While faster for somethings....that's a pretty large set of caveats!
I have a issue to improve performance where I can push this forward: https://github.com/davesnx/query-json/issues/7
But sure, are caveats!
https://github.com/davesnx/query-json#performance https://github.com/davesnx/query-json/blob/master/benchmarks...
But all explanations aren't based by any evidence, just asumptions.
IMO, an individual dev making a fast useful tool should always be welcomed as a feat of worthy hacking.
But if I had done something like that, and then serendipitously discovered that I was exceeding the original's performance, I certainly wouldn't be shy about it.
Also, this comes across as armchair criticism purely for the sake of armchair criticism. My own experience has been that, when I'm doing ETL that involves wrangling JSON, the "wrangling JSON" bit of it is almost always the bottleneck. So any improvement is more than welcome and deserves to be cheered. Even if it's an improvement on something that's already the current fastest way to do it.
Reimplementing a piece of software that is 12 years old which mimics their UX and improves performance and error messages it's more than welcome in my opinion. My purpose was to learn the OCaml stack of writting compilers, so I personally found that I "needed" a language already created.
Thanks for raising those concerns
We'd all be better off if plebes grew their skills by reimplementing common tools.
I used Reason to compile to Native, so using OCaml's stdlib and OCaml's dependencies and compiling it with OCaml, but my source code is written in Reason syntax.
In this case the author is using Reason as an alternative syntax to OCaml. Reason resembles javascript a little more, and some people find that nicer to work with. So the idea is that you write Reason code, then translate it into OCaml code using the Reason tools, and then ultimately you compile it down to a native binary.
If instead you want to write a web-app which runs in a web browser or node.js, then you'd need to compile it to Javascript, which is what bucklescript helps you do.
Where does Rescript come in? As explained above, Reason can be used for writing either native apps or javascript apps. However, it's hard to evolve the syntax of Reason in a way which satisfies both aims. So they've now split the work -- going forward, Reason will specialize on native, and Rescript will specialize on javascript apps. Their syntax is expected to diverge from each other, in order to support those aims as best as they can.
It wasn't.
Before doing any curl | bash, check what's on the install command, that's the entire point of it.
I'm sure other people have use cases where the browser wouldn't meet their needs, but for me, I find jq unnecessary.
I was somewhat surprised it didn't use an existing json parser library.
I often have to pluck out attributes from streams of json records (1 json object per line) - often millions/billions.
jq is almost always the bottleneck in the pipeline at 100% CPU - so much so that we often add an fgrep to the left side of the pipeline to minimize the input to jq as much as possible.
The main weakness seems to be streaming use cases (not having the whole file in memory at once). These are supported, but the syntax is quite awkward.
I should specify that on the performance section, Thanks!
The C code is clean enough as C code goes, but fairly monolithic. And it’s C, so it’s not noticeably slow until you start processing GB. But it would probably take a rewrite to improve its performance significantly.
https://mosermichael.github.io/jq-illustrated/dir/content.ht...
and there are a few techniques to "discover" the schema of the json file, I trend to read with '.' or 'keys' and later keep going.
I'm planning to implement a flag where each operation prints the internal state of the json, so you would see what are the "pipes".
I will pick a few of your cheatsheet to implement next in q, Thanks!
It is a JSON parser in C without heap allocations. The query language is piddly, but the tool can be useful for grabbing a single value from a very large JSON file. I don't have time for it, but someone could fork and make it a real deal.
Menhir/sedlex and others are pretty high accessibility barrier for new commers.
One of the nice things about all of it it's the discord, it's friendly and always helpful.
Hope it helps, just let me know if there's any specific!
For the set of operations that I implement it it's faster, that's true.
Replacing that with, say, traditional command line flags would make it a lot less useful for me, I'd probably have to build much longer pipe-chains to do things that are relatively simple and readable jq snippets (if one knows the syntax.)
Using an established scripting language in its place would make it pretty much just python -c/ruby -e or whatever with some pre-loaded functions, but what's the point? You can always just write a quick python/ruby/whatever script, jq to me is an alternative for cases where a script feels unnecessary. It would also mean everything gets more verbose, so less of my jq transformations can be inlined without loss of readability.
Aligning it to more established languages would probably cause confusion as well in those cases where it doesn't match the reference language 1:1. Looks like javascript, writes like javascript, but only for a tiny subset of the language, etc.
Doing this only for a few function names or syntax constructs still results in a pretty unique and unusual language that will require people to reference the docs a lot, just now lots of existing scripts break.
There're a lot of quirks from the usage of it and people struggling with learning such a great tool, so in the area of query-json it will try to make a better interface for users.
The feature that I think penalizes a lot jq is "def functions", the capacity of define any function that can be available during run-time.
This creates a few layers, one of the difference is the interpreter and the linker, the responsible for getting all the builtin functions and compile them have them ready to use at runtime.
The other pain point is the architecture of the operations on top of jq, since it's a stack based. In query-json it's a piped recursive operations.
Aside from the code, the OCaml stack, menhir has been proved to be really fast when creating those kind of compilers.
I will dig more into performance and try to profile both tools in order to improve mine.
Thanks
For example, someone who works with a lot of Python might prefer something like https://github.com/kellyjonbrazil/jello to write comprehensions using the full capabilities of Python, especially since that would provide a direct path to using the final expressions in a Python program or even embedded in one of the environments where Python is used as a scripting language. Is that a viable alternative? The answer depends entirely on who's asking.